22,027 Darija multiple-choice questions across 44 subjects, translated from selected MMLU and ArabicMMLU subsets.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | Moroccan Darija multiple-choice exam QA translated from MMLU and ArabicMMLU |
| Page status | active |
| Metric | accuracy (size-weighted mean across tasks) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 22027 |
| Dataset licence | MIT (Hub card, linking hendrycks/test); ArabicMMLU parent licence is stated two ways on arabic_mmlu.md (GitHub CC-BY-NC-SA-4.0 vs Hub cc-by-nc-4.0) |
| Publisher | MBZUAI-Paris, with EMINES-UM6P, LINAGORA, KTH, AtlasIA and École Polytechnique |
DarijaMMLU tests whether a model can answer subject questions written in Moroccan Darija. Forty-four Hub configs mix two sources: selected English [MMLU](mmlu.md) subjects and selected [ArabicMMLU](arabic_mmlu.md) subjects, both translated with Claude 3.5 Sonnet. Items have two to five options. Some ArabicMMLU-derived rows carry a context passage. This is not a native Darija exam corpus and not a full copy of either parent benchmark.
Multiple-choice QA in Darija. lm-eval group darijammlu prompts with a Darija instruction, the subject name, the question, and A.–E. options, then scores letter accuracy. Template: test_split test, fewshot_split dev, sampler first_n. YAML does not pin num_fewshot; the Atlas-Chat table reports 0-shot and 3-shot.
No model card in ModelSpec reports this benchmark yet.