983 Arabic multiple-choice questions on Arabic language and grammar, an exact two-subject slice of ArabicMMLU repackaged standalone; now tracked in the OALL v2 Arabic leaderboard.
unassessed
| Category | domain |
|---|---|
| Subcategory | Arabic-language and Arabic-grammar multiple-choice questions, a two-subject slice of ArabicMMLU redistributed as its own dataset |
| Page status | active |
| Metric | exact_match (accuracy on the selected option) in HELM; acc_norm (normalized accuracy) in the OALL v2 / LightEval implementation |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 983 |
| Dataset licence | CC BY-NC 4.0, per the Hugging Face dataset card; the parent ArabicMMLU project's own paper separately states a CC BY 4.0 licence, so this specific redistribution is more restrictive (noncommercial) than the source project's own stated licence. |
| Publisher | Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), with Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi and the University of Melbourne (ArabicMMLU co-authors' affiliations) |
MadinahQA tests knowledge of Arabic language and grammar through two multiple-choice subject sets, "Arabic Language (General)" and "Arabic Language (Grammar)". Both are drawn, unchanged in row count, from ArabicMMLU (Koto et al., 2024), a 40-subject, 14,575-question Arabic knowledge benchmark built from real school, university and professional exam questions across North Africa, the Levant and the Gulf: ArabicMMLU's own paper reports exactly 615 questions for "Arabic Language (General)" and 368 for "Arabic Language (Grammar)", matching this dataset's configs exactly. MBZUAI, the same publisher behind ArabicMMLU, released these two subjects as their own standalone Hugging Face dataset under the name MadinahQA; the origin of that specific name (for example, a particular exam board or source) is not established from the sources reviewed for this page. The General subject includes a reading-passage Context field for most (602 of 612) of its test questions; the Grammar subject has none.
An Arabic-language multiple-choice question (with a reading passage for most "General" items), offering up to five options, in; the model selects the correct option. HELM evaluates it as a joint multiple-choice task with an Arabic-language instruction ("السؤال التالي هو سؤال متعدد الإختيارات. اختر الإجابة الصحيحة" -- "the following is a multiple-choice question, choose the correct answer") and Arabic reference letters (أ ب ج د هـ) rather than Latin A-E.
No model card in ModelSpec reports this benchmark yet.