HELM's Arabic Enterprise finance set: 299 textbook-derived items in three formats (3-way MCQ, yes/no, numeric calculation), in Arabic or English.
unassessed
| Category | domain |
|---|---|
| Subcategory | Arabic (and English) finance textbook QA: 3-way MCQ, yes/no, numeric calculation |
| Page status | unknown |
| Metric | exact_match (MCQ); quasi_exact_match (bool, schema headline); calculation_accuracy (calculation) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 299 |
| Dataset licence | CC-BY-4.0 |
| Publisher | Stanford CRFM (HELM Arabic Enterprise) |
arabic_finance is HELM's finance slice of the stanford-crfm/arabic-enterprise dataset. Each item is a short finance question drawn, per HELM's schema, from English-language finance textbooks and machine-translated into Arabic. The same 299 rows are stored with English and Arabic question, choice and answer fields. HELM splits them into three task formats: 23 three-option multiple-choice questions (task=mcq), 57 yes/no verifications (task=bool), and 219 numeric calculation problems (task=calcu). A run is one format and one language (default Arabic). It is text-only professional-finance QA, not [financebench](financebench.md) (English SEC-filing QA) and not [financeiq](financeiq.md).
Three HELM run specs share one CSV. MCQ: pick one of three labelled choices (CSV uses A/B/C; the Arabic adapter remaps prefixes to أ/ب/ج). Bool: answer نعم/لا or Yes/No. Calculation: write reasoning, then a numeric answer inside \\boxed{}; an LLM annotator judges mathematical equivalence.
No model card in ModelSpec reports this benchmark yet.