Open Arabic LLM Leaderboard — Complete configuration

An lm-evaluation-harness reproduction of the original Open Arabic LLM Leaderboard: 14 native and machine-translated Arabic task groups aggregated into one size-weighted accuracy score.

Also known as: OALL, Open Arabic LLM Leaderboard, OALL v1

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryleaderboard aggregating 14 Arabic-language task groups (native and machine-translated) into one weighted score
Page statussuperseded
Metricaccuracy (acc and acc_norm, size-weighted mean across 14 task groups)
Directionhigher_is_better
Unit%
Dataset licenceVaries by component dataset (AlGhafa, ACVA/AceGPT, EXAMS and each translated source benchmark carry their own licence); not established as a single licence for this page.
PublisherOpen Arabic LLM Leaderboard (OALL) project: Technology Innovation Institute (TII) and 2A2I, with Hugging Face

What it measures

arabic_leaderboard_complete is not one benchmark but a leaderboard configuration: an lm-evaluation-harness task group, named to match the Open Arabic LLM Leaderboard (OALL), that runs 14 separately-authored Arabic evaluation task groups and aggregates their scores into one weighted number. The 14 groups mix native-Arabic content -- AlGhafa (reading comprehension, sentiment analysis and question answering, built by the TII LLM team, part natively authored and part translated), ACVA (58 yes/no topic datasets on Arabic culture and values, generated by GPT-3.5 Turbo), and Arabic EXAMS (translated high-school exam questions) -- with ten "arabic_mt_*" groups that are machine translations of well-known English benchmarks: ARC-Challenge, ARC-Easy, BoolQ, COPA, HellaSwag, MMLU, OpenBookQA, PIQA, RACE, SciQ and ToxiGen. A model's headline score on this configuration says less about any single skill than about a broad, unevenly-weighted mix of Arabic knowledge, commonsense reasoning, reading comprehension and toxicity classification.

Task format

A group-of-groups: each of the 14 component tasks is itself multiple-choice or yes/no question answering (several, such as AlGhafa and Arabic MMLU, are themselves further subdivided into many per-topic or per-subject tasks). No single prompt format applies across the whole aggregation.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub