AlGhafa

TII's Arabic multiple-choice suite; HELM and OALL score nine native Hugging Face configs with public answers, not the translated COPA/OpenBookQA extras.

Also known as: AlGhafa Evaluation Benchmark, AlGhafa-Arabic-LLM-Benchmark-Native, arabic_leaderboard_alghafa

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategoryArabic multiple-choice suite (reading, exams, sentiment, dialect identification)
Page statusactive
Metricexact_match (HELM); acc / acc_norm (lm-evaluation-harness OALL group)
Directionhigher_is_better
Unit%
Dataset size22932
PublisherTechnology Innovation Institute (TII); dataset hosted by OALL

What it measures

AlGhafa is a multiple-choice evaluation for Arabic language models. A model sees an Arabic query and labelled options, then must pick the correct option. The native packaging mixes reading (Belebele MSA and dialects, SOQAL, XGLUE-MLQA), school exams, AraFacts true/false, hotel-review sentiment, and Twitter sentiment. It is Arabic text, not [arabic_mmlu](arabic_mmlu.md) (natively sourced school exams by subject) and not the English originals those reading sets come from.

Task format

Multiple-choice with two to five options depending on the subset (true/false, 3-way sentiment, 4-way exams/Belebele, 5-way SOQAL/XGLUE). HELM uses joint generation with Arabic letter prefixes (أ–هـ) and exact_match. lm-evaluation-harness OALL configs score log-likelihood accuracy (acc and acc_norm) on the same Native configs.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub