A 29-language, expert-verified translation of MMLU-Pro with 11,829 identical questions per language, built to compare cross-linguistic reasoning rather than just English knowledge.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | multilingual multitask academic and professional knowledge (translated MMLU-Pro, 29 languages) |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 342441 |
| Dataset licence | MIT |
| Publisher | The University of Tokyo, with co-authors across 16 further institutions including Duke-NUS Medical School, Waseda University, Carnegie Mellon University and Yale University |
MMLU-ProX takes MMLU-Pro's ten-option, reasoning-heavy multiple-choice questions across 14 categories and translates every question into 29 typologically diverse languages, keeping the same 11,829 questions identical (parallel) across every language version so scores are directly comparable across languages rather than only within one. Translation runs through multiple large language models followed by expert human review: over 30 professional translators, native in the target language and proficient in English, rated accuracy, fluency and completeness on a 1-5 scale for a stratified sample, triggering full retranslation of any subject-language pair that scored below 3 on any dimension (only Yoruba law needed this). A separate lite version keeps 658 questions per language (a fixed 47-question-per-category-ish, difficulty-consistent subsample) for cheaper evaluation runs.
Ten-option (occasionally fewer) multiple-choice questions across the same 14 MMLU-Pro categories (Biology, Business, Chemistry, Computer Science, Economics, Engineering, Health, History, Law, Math, Philosophy, Physics, Psychology, Other), graded on the single option selected, in each of 29 languages. Following MMLU-Pro's own protocol, the paper's primary reported setting is 5-shot chain-of-thought prompting, though the authors also evaluate 0-shot for comparison across their 36-model sweep.
No model card in ModelSpec reports this benchmark yet.