EgyMMLU

Egyptian Arabic multiple-choice exam of 22,027 items across 44 subjects, translated from MMLU and ArabicMMLU.

Also known as: Egy MMLU, UBC-NLP/EgyMMLU

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryEgyptian Arabic multitask multiple-choice (44 subjects translated from MMLU and ArabicMMLU)
Page statusactive
Metricaccuracy (acc), size-weighted across subjects in group egymmlu
Directionhigher_is_better
Unit%
Dataset size22027
Dataset licenceMIT
PublisherUBC-NLP (University of British Columbia)

What it measures

EgyMMLU tests whether a model can answer multiple-choice questions written in Egyptian Arabic. The 44 subjects mix English MMLU topics (for example professional_law, moral_scenarios) with ArabicMMLU topics (islamic_studies, driving_test, arabic_language). Some items include a context passage. The skill is dialectal exam QA, not a from-scratch Egyptian curriculum. It is not native [ArabicMMLU](arabic_mmlu.md) and not English [MMLU](mmlu.md).

Task format

Multiple-choice with lettered options A-E as needed. lm-eval builds an Egyptian-Arabic prompt from egy_subject, optional context, question, and choices. Target is the integer answer index mapped to a letter. Group egymmlu aggregates size-weighted acc across subjects.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub