Egyptian Arabic multiple-choice exam of 22,027 items across 44 subjects, translated from MMLU and ArabicMMLU.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | Egyptian Arabic multitask multiple-choice (44 subjects translated from MMLU and ArabicMMLU) |
| Page status | active |
| Metric | accuracy (acc), size-weighted across subjects in group egymmlu |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 22027 |
| Dataset licence | MIT |
| Publisher | UBC-NLP (University of British Columbia) |
EgyMMLU tests whether a model can answer multiple-choice questions written in Egyptian Arabic. The 44 subjects mix English MMLU topics (for example professional_law, moral_scenarios) with ArabicMMLU topics (islamic_studies, driving_test, arabic_language). Some items include a context passage. The skill is dialectal exam QA, not a from-scratch Egyptian curriculum. It is not native [ArabicMMLU](arabic_mmlu.md) and not English [MMLU](mmlu.md).
Multiple-choice with lettered options A-E as needed. lm-eval builds an Egyptian-Arabic prompt from egy_subject, optional context, question, and choices. Target is the integer answer index mapped to a letter. Group egymmlu aggregates size-weighted acc across subjects.
No model card in ModelSpec reports this benchmark yet.