MMLU's test set professionally translated into 14 languages, testing whether a model's broad academic knowledge holds up outside English.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | multilingual knowledge and reasoning |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 196588 |
| Dataset licence | MIT |
| Publisher | OpenAI |
MMMLU takes the original MMLU test questions - four-choice questions across 57 subjects spanning STEM, humanities, social sciences and other professional topics - and evaluates the same questions in 14 non-English languages, using human (not machine) translation. It measures whether a model's knowledge and reasoning ability, as captured by MMLU in English, transfers to other languages, including lower-resource ones such as Yoruba and Swahili, rather than measuring translation quality itself.
Four-option multiple-choice question answering, one language per evaluation run, typically zero-shot or few-shot.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 92.67 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 91.1 | 2026-04 |