Arabic Legal (HELM Arabic Enterprise)

HELM's Arabic Enterprise legal set: 200 UAE-law questions scored by an LLM judge, in closed-book and open-book (statute-in-prompt) modes.

Also known as: arabic_legal_qa, arabic_legal_rag, Arabic Enterprise legal

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryUAE-law Arabic short-answer QA, closed-book and open-book
Page statusunknown
Metricmodel_judged_score
Directionhigher_is_better
Unit%
Dataset size200
Dataset licenceCC-BY-4.0
PublisherStanford CRFM (HELM Arabic Enterprise)

What it measures

arabic_legal is HELM's legal slice of stanford-crfm/arabic-enterprise. Each item is an open-ended question in Arabic about United Arab Emirates law, paired with a short Arabic reference answer and a statute-like context passage. HELM's schema says Arabic legal experts wrote the questions. Two protocols share the same 200 rows: closed-book QA sends only the question; open-book RAG prepends the context, then a blank line, then the question. The adapter instruction in both cases tells the model to answer briefly in Modern Standard Arabic in the setting of UAE law, and to emit the answer only. This is not [legalbench](legalbench.md) (English lawyer-authored classification tasks) and not [arabic_exams](arabic_exams.md).

Task format

Short-answer generation in Arabic. arabic_legal_qa: question only. arabic_legal_rag: context plus question. An LLM annotator scores 1 if the output is equivalent to the reference answer, else 0.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub