HELM's Arabic Enterprise legal set: 200 UAE-law questions scored by an LLM judge, in closed-book and open-book (statute-in-prompt) modes.
unassessed
| Category | domain |
|---|---|
| Subcategory | UAE-law Arabic short-answer QA, closed-book and open-book |
| Page status | unknown |
| Metric | model_judged_score |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 200 |
| Dataset licence | CC-BY-4.0 |
| Publisher | Stanford CRFM (HELM Arabic Enterprise) |
arabic_legal is HELM's legal slice of stanford-crfm/arabic-enterprise. Each item is an open-ended question in Arabic about United Arab Emirates law, paired with a short Arabic reference answer and a statute-like context passage. HELM's schema says Arabic legal experts wrote the questions. Two protocols share the same 200 rows: closed-book QA sends only the question; open-book RAG prepends the context, then a blank line, then the question. The adapter instruction in both cases tells the model to answer briefly in Modern Standard Arabic in the setting of UAE law, and to emit the answer only. This is not [legalbench](legalbench.md) (English lawyer-authored classification tasks) and not [arabic_exams](arabic_exams.md).
Short-answer generation in Arabic. arabic_legal_qa: question only. arabic_legal_rag: context plus question. An LLM annotator scores 1 if the output is equivalent to the reference answer, else 0.
No model card in ModelSpec reports this benchmark yet.