Open Arabic LLM Leaderboard — Light configuration

A 10%-random-sample version of arabic_leaderboard_complete's same 14 Arabic task groups, run for lower cost; ACVA's 10-item Yemen subset is kept at full size rather than sampled further.

Also known as: OALL Light

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategory10%-sample version of arabic_leaderboard_complete's 14 task groups
Page statussuperseded
Metricaccuracy (acc and acc_norm, size-weighted mean across 14 sampled task groups)
Directionhigher_is_better
Unit%
Dataset licenceVaries by component dataset; see arabic_leaderboard_complete.
PublisherOpen Arabic LLM Leaderboard (OALL) project: Technology Innovation Institute (TII) and 2A2I, with Hugging Face

What it measures

arabic_leaderboard_light runs the same 14 component task groups as arabic_leaderboard_complete -- AlGhafa, ACVA, Arabic EXAMS, and ten machine-translated groups (ARC-Challenge, ARC-Easy, BoolQ, COPA, HellaSwag, MMLU, OpenBookQA, PIQA, RACE, SciQ, ToxiGen) -- but scores each one against a random 10% sample of its test set rather than the full set, to cut evaluation cost. The harness's own README for this task documents one explicit exception: ACVA's Yemen subset has only 10 test items to begin with, so it is evaluated at full size rather than reduced to a single item.

Task format

Identical task formats to arabic_leaderboard_complete's 14 components, each applied to a smaller, randomly sampled item set.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub