Multilingual MMLU (MMMLU)

MMLU's test set professionally translated into 14 languages, testing whether a model's broad academic knowledge holds up outside English.

Also known as: MMMLU, Multilingual MMLU

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorymultilingual knowledge and reasoning
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size196588
Dataset licenceMIT
PublisherOpenAI

What it measures

MMMLU takes the original MMLU test questions - four-choice questions across 57 subjects spanning STEM, humanities, social sciences and other professional topics - and evaluates the same questions in 14 non-English languages, using human (not machine) translation. It measures whether a model's knowledge and reasoning ability, as captured by MMLU in English, transfers to other languages, including lower-resource ones such as Yoruba and Swahili, rather than measuring translation quality itself.

Task format

Four-option multiple-choice question answering, one language per evaluation run, typically zero-shot or few-shot.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic92.672026-04
Claude Opus 4.6Anthropic91.12026-04

Data

This page as JSON · Edit on GitHub