DarijaMMLU

22,027 Darija multiple-choice questions across 44 subjects, translated from selected MMLU and ArabicMMLU subsets.

Also known as: Darija MMLU, MBZUAI-Paris/DarijaMMLU

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryMoroccan Darija multiple-choice exam QA translated from MMLU and ArabicMMLU
Page statusactive
Metricaccuracy (size-weighted mean across tasks)
Directionhigher_is_better
Unit%
Dataset size22027
Dataset licenceMIT (Hub card, linking hendrycks/test); ArabicMMLU parent licence is stated two ways on arabic_mmlu.md (GitHub CC-BY-NC-SA-4.0 vs Hub cc-by-nc-4.0)
PublisherMBZUAI-Paris, with EMINES-UM6P, LINAGORA, KTH, AtlasIA and École Polytechnique

What it measures

DarijaMMLU tests whether a model can answer subject questions written in Moroccan Darija. Forty-four Hub configs mix two sources: selected English [MMLU](mmlu.md) subjects and selected [ArabicMMLU](arabic_mmlu.md) subjects, both translated with Claude 3.5 Sonnet. Items have two to five options. Some ArabicMMLU-derived rows carry a context passage. This is not a native Darija exam corpus and not a full copy of either parent benchmark.

Task format

Multiple-choice QA in Darija. lm-eval group darijammlu prompts with a Darija instruction, the subject name, the question, and A.–E. options, then scores letter accuracy. Template: test_split test, fewshot_split dev, sampler first_n. YAML does not pin num_fewshot; the Atlas-Chat table reports 0-shot and 3-shot.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub