A 10%-random-sample version of arabic_leaderboard_complete's same 14 Arabic task groups, run for lower cost; ACVA's 10-item Yemen subset is kept at full size rather than sampled further.
unassessed
| Category | composite |
|---|---|
| Subcategory | 10%-sample version of arabic_leaderboard_complete's 14 task groups |
| Page status | superseded |
| Metric | accuracy (acc and acc_norm, size-weighted mean across 14 sampled task groups) |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | Varies by component dataset; see arabic_leaderboard_complete. |
| Publisher | Open Arabic LLM Leaderboard (OALL) project: Technology Innovation Institute (TII) and 2A2I, with Hugging Face |
arabic_leaderboard_light runs the same 14 component task groups as arabic_leaderboard_complete -- AlGhafa, ACVA, Arabic EXAMS, and ten machine-translated groups (ARC-Challenge, ARC-Easy, BoolQ, COPA, HellaSwag, MMLU, OpenBookQA, PIQA, RACE, SciQ, ToxiGen) -- but scores each one against a random 10% sample of its test set rather than the full set, to cut evaluation cost. The harness's own README for this task documents one explicit exception: ACVA's Yemen subset has only 10 test items to begin with, so it is evaluated at full size rather than reduced to a single item.
Identical task formats to arabic_leaderboard_complete's 14 components, each applied to a smaller, randomly sampled item set.
No model card in ModelSpec reports this benchmark yet.