An lm-evaluation-harness reproduction of the original Open Arabic LLM Leaderboard: 14 native and machine-translated Arabic task groups aggregated into one size-weighted accuracy score.
unassessed
| Category | composite |
|---|---|
| Subcategory | leaderboard aggregating 14 Arabic-language task groups (native and machine-translated) into one weighted score |
| Page status | superseded |
| Metric | accuracy (acc and acc_norm, size-weighted mean across 14 task groups) |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | Varies by component dataset (AlGhafa, ACVA/AceGPT, EXAMS and each translated source benchmark carry their own licence); not established as a single licence for this page. |
| Publisher | Open Arabic LLM Leaderboard (OALL) project: Technology Innovation Institute (TII) and 2A2I, with Hugging Face |
arabic_leaderboard_complete is not one benchmark but a leaderboard configuration: an lm-evaluation-harness task group, named to match the Open Arabic LLM Leaderboard (OALL), that runs 14 separately-authored Arabic evaluation task groups and aggregates their scores into one weighted number. The 14 groups mix native-Arabic content -- AlGhafa (reading comprehension, sentiment analysis and question answering, built by the TII LLM team, part natively authored and part translated), ACVA (58 yes/no topic datasets on Arabic culture and values, generated by GPT-3.5 Turbo), and Arabic EXAMS (translated high-school exam questions) -- with ten "arabic_mt_*" groups that are machine translations of well-known English benchmarks: ARC-Challenge, ARC-Easy, BoolQ, COPA, HellaSwag, MMLU, OpenBookQA, PIQA, RACE, SciQ and ToxiGen. A model's headline score on this configuration says less about any single skill than about a broad, unevenly-weighted mix of Arabic knowledge, commonsense reasoning, reading comprehension and toxicity classification.
A group-of-groups: each of the 14 component tasks is itself multiple-choice or yes/no question answering (several, such as AlGhafa and Arabic MMLU, are themselves further subdivided into many per-topic or per-subject tasks). No single prompt format applies across the whole aggregation.
No model card in ModelSpec reports this benchmark yet.