A harder, ten-option successor to MMLU with about 12,000 reasoning-heavy questions across 14 categories, built to restore headroom lost to MMLU's saturation at the frontier.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | multitask academic and professional knowledge (reasoning-augmented) |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 12032 |
| Dataset licence | MIT |
| Publisher | TIGER Lab, University of Waterloo (with co-authors at the University of Toronto and Carnegie Mellon University) |
MMLU-Pro gives a model a question and up to ten labelled answer options (most questions use all ten; a small number carry fewer after manual review removed unreasonable distractors), drawn from 14 broad categories spanning STEM, humanities, social sciences and business/health/other topics: Biology, Business, Chemistry, Computer Science, Economics, Engineering, Health, History, Law, Math, Philosophy, Physics, Psychology and Other. The model must select the single correct option. Roughly 57% of the questions are a difficulty-filtered subset of the original MMLU test set (with "trivial and ambiguous" items removed); the remaining 43% are newly written from a STEM question website, TheoremQA and SciBench, with GPT-4-generated distractors expanding every question from four options toward ten, reviewed afterward by a panel of over ten subject-matter experts.
Multiple-choice question answering with up to ten labelled options, graded on the single option the model selects. The reference protocol is 5-shot with chain-of-thought prompting, with fewshot examples drawn from a dedicated 70-question validation split; scoring extracts the answer letter from free-form generated text (for example via a regex matching "the answer is (X)") rather than comparing option log-likelihoods, because the authors found direct/log-likelihood scoring under-performs chain-of-thought by up to 19 points on this dataset -- the opposite of the original MMLU. Across 24 prompt styles the authors tested, score sensitivity to prompt wording fell from 4-5% on MMLU to about 2% on MMLU-Pro.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| GPT-5.1 | OpenAI | 86.0 | 2026-04 |
| GPT-5.1 Chat | OpenAI | 86.0 | 2026-04 |
| GPT-5.1 Codex | OpenAI | 86.0 | 2026-04 |
| GPT-5.1 Codex Max | OpenAI | 86.0 | 2026-04 |
| GPT-5.1 Codex mini | OpenAI | 86.0 | 2026-04 |
| o3-pro | OpenAI | 85.5 | 2026-04 |
| Claude Mythos Preview | Anthropic | 85.2 | 2026-04 |
| Claude Opus 4 | Anthropic | 85.2 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 85.2 | 2026-04 |
| GPT-5 | OpenAI | 84.8 | 2026-04 |
| GPT-5 Chat (latest) | OpenAI | 84.8 | 2026-04 |
| GPT-5 Pro | OpenAI | 84.8 | 2026-04 |
| GPT-5-Codex | OpenAI | 84.8 | 2026-04 |
| GPT-5.2 | OpenAI | 84.8 | 2026-04 |
| GPT-5.2 Chat | OpenAI | 84.8 | 2026-04 |
| GPT-5.2 Codex | OpenAI | 84.8 | 2026-04 |
| GPT-5.2 Pro | OpenAI | 84.8 | 2026-04 |
| GPT-5.3 Chat (latest) | OpenAI | 84.8 | 2026-04 |
| GPT-5.3 Codex | OpenAI | 84.8 | 2026-04 |
| GPT-5.3 Codex Spark | OpenAI | 84.8 | 2026-04 |
| GPT-5.4 | OpenAI | 84.8 | 2026-04 |
| GPT-5.4 mini | OpenAI | 84.8 | 2026-04 |
| GPT-5.4 nano | OpenAI | 84.8 | 2026-04 |
| GPT-5.4 Pro | OpenAI | 84.8 | 2026-04 |
| Claude Opus 4.5 | Anthropic | 84.5 | 2026-04 |
| Claude Opus 4.5 (latest) | Anthropic | 84.5 | 2026-04 |
| o3 | OpenAI | 84.1 | 2026-04 |
| o3-deep-research | OpenAI | 84.1 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 84.0 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 84.0 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 84.0 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 84.0 | 2026-04 |
| Claude Opus 4.1 | Anthropic | 83.8 | 2026-04 |
| Claude Opus 4.1 (latest) | Anthropic | 83.8 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 83.0 | 2026-04 |
| Claude Sonnet 4.6 | Anthropic | 83.0 | 2026-04 |
| Grok 4 | xAI | 82.5 | 2026-04 |
| Grok 4 Fast | xAI | 82.5 | 2026-04 |
| Grok 4 Fast (Non-Reasoning) | xAI | 82.5 | 2026-04 |
| Grok 4.1 Fast | xAI | 82.5 | 2026-04 |
| Grok 4.1 Fast (Non-Reasoning) | xAI | 82.5 | 2026-04 |
| Grok 4.20 (Non-Reasoning) | xAI | 82.5 | 2026-04 |
| Grok 4.20 (Reasoning) | xAI | 82.5 | 2026-04 |
| Grok 4.20 Multi-Agent | xAI | 82.5 | 2026-04 |
| Claude Opus 4 (latest) | Anthropic | 82.1 | 2026-04 |
| o1-pro | OpenAI | 82.0 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 81.7 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 81.7 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 81.0 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 81.0 | 2026-04 |
| o4-mini | OpenAI | 81.0 | 2026-04 |
| o4-mini-deep-research | OpenAI | 81.0 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 80.2 | 2026-04 |
| GPT-4.1 | OpenAI | 80.1 | 2026-04 |
| DeepSeek R1 | DeepSeek | 79.8 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 79.8 | 2026-04 |
| GPT-5 Mini | OpenAI | 79.5 | 2026-04 |
| o3-mini | OpenAI | 79.0 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 78.8 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 78.8 | 2026-04 |
| Claude Sonnet 4 (latest) | Anthropic | 78.5 | 2026-04 |
| Qwen3-Coder 480B-A35B Instruct | Alibaba / Qwen Team | 78.5 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Image | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Image (Preview) | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Lite | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 06-17 | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 09-25 | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Preview 04-17 | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Preview 05-20 | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Preview 09-25 | Google DeepMind | 78.2 | 2026-04 |
| Gemini 2.5 Flash Preview TTS | Google DeepMind | 78.2 | 2026-04 |
| Claude Sonnet 3.7 | Anthropic | 78.0 | 2026-04 |
| o1 | OpenAI | 78.0 | 2026-04 |
| o1-preview | OpenAI | 78.0 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 77.2 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 76.2 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 76.2 | 2026-04 |
| Grok 3 | xAI | 76.0 | 2026-04 |
| Grok 3 Fast | xAI | 76.0 | 2026-04 |
| Grok 3 Fast Latest | xAI | 76.0 | 2026-04 |
| Grok 3 Latest | xAI | 76.0 | 2026-04 |
| DeepSeek Chat | DeepSeek | 75.5 | 2026-04 |
| DeepSeek V2 | DeepSeek | 75.5 | 2026-04 |
| DeepSeek V2 Lite | DeepSeek | 75.5 | 2026-04 |
| DeepSeek V2 Lite Chat | DeepSeek | 75.5 | 2026-04 |
| DeepSeek V3 | DeepSeek | 75.5 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 75.5 | 2026-04 |
| Claude Haiku 4.5 | Anthropic | 75.2 | 2026-04 |
| Claude Haiku 4.5 (latest) | Anthropic | 75.2 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 74.8 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 74.8 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 74.8 | 2026-04 |
| GPT-4.1 mini | OpenAI | 74.2 | 2026-04 |
| Magistral Medium (latest) | Mistral AI | 74.2 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 74.1 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 74.1 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 74.1 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 74.1 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 73.5 | 2026-04 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 73.5 | 2026-04 |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | Meta | 73.5 | 2026-04 |
| GPT-4o | OpenAI | 72.6 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 72.6 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 72.6 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 72.6 | 2026-04 |
| o1-mini | OpenAI | 72.5 | 2026-04 |
| QwQ Plus | Alibaba / Qwen Team | 72.5 | 2026-04 |
| Gemma 4 26B | Google DeepMind | 72.3 | 2026-04 |
| gemma 4 26B A4B it | Google DeepMind | 72.3 | 2026-04 |
| gemma 4 26B A4B it GGUF | Unsloth | 72.3 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 72.0 | 2026-04 |
| Qwen2.5 72B Instruct | Alibaba / Qwen Team | 71.5 | 2026-04 |
| GPT-5 Nano | OpenAI | 70.3 | 2026-04 |
| Grok 3 Mini | xAI | 70.2 | 2026-04 |
| Grok 3 Mini Fast | xAI | 70.2 | 2026-04 |
| Grok 3 Mini Fast Latest | xAI | 70.2 | 2026-04 |
| Grok 3 Mini Latest | xAI | 70.2 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 69.8 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 69.8 | 2026-04 |
| Mistral Large 3 | Mistral AI | 69.8 | 2026-04 |
| Claude Opus 3 | Anthropic | 68.5 | 2026-04 |
| phi 4 | Microsoft | 68.5 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 68.5 | 2026-04 |
| Qwen3 30B A3B Instruct 2507 | Alibaba / Qwen Team | 68.5 | 2026-04 |
| Qwen3 30B A3B NVFP4 | NVIDIA | 68.5 | 2026-04 |
| Qwen3 30B-A3B | Alibaba / Qwen Team | 68.5 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 68.3 | 2026-04 |
| Llama 4 Scout 17B 16E | Meta | 68.2 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 68.2 | 2026-04 |
| Llama-4-Scout-17B-16E-Instruct-FP8 | Meta | 68.2 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 67.5 | 2026-04 |
| Llama 3.1 405B | Meta | 67.5 | 2026-04 |
| Llama 3.1 405B FP8 | Meta | 67.5 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 67.5 | 2026-04 |
| Llama 3.1 405B Instruct FP8 | Meta | 67.5 | 2026-04 |
| GPT-4 Turbo | OpenAI | 67.4 | 2026-04 |
| GPT-4o mini | OpenAI | 66.5 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 66.5 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 66.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 65.5 | 2026-04 |
| Grok 2 | xAI | 65.5 | 2026-04 |
| Grok 2 (1212) | xAI | 65.5 | 2026-04 |
| Grok 2 Latest | xAI | 65.5 | 2026-04 |
| Grok 2 Vision | xAI | 65.5 | 2026-04 |
| Grok 2 Vision (1212) | xAI | 65.5 | 2026-04 |
| Grok 2 Vision Latest | xAI | 65.5 | 2026-04 |
| Claude Haiku 3.5 | Anthropic | 65.3 | 2026-04 |
| Claude Haiku 3.5 (latest) | Anthropic | 65.3 | 2026-04 |
| Qwen2.5 32B Instruct | Alibaba / Qwen Team | 65.2 | 2026-04 |
| Qwen2.5 32B Instruct AWQ | Alibaba / Qwen Team | 65.2 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 64.8 | 2026-04 |
| Magistral Small | Mistral AI | 64.5 | 2026-04 |
| Magistral Small 2506 | Mistral AI | 64.5 | 2026-04 |
| Qwen3 14B | Alibaba / Qwen Team | 64.2 | 2026-04 |
| Qwen3 14B AWQ | Alibaba / Qwen Team | 64.2 | 2026-04 |
| Qwen3 14B NVFP4 | NVIDIA | 64.2 | 2026-04 |
| GPT-4.1 nano | OpenAI | 64.1 | 2026-04 |
| Command A | Cohere | 63.8 | 2026-04 |
| Command A Reasoning | Cohere | 63.8 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 63.5 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 63.5 | 2026-04 |
| Llama 3.1 70B | Meta | 60.8 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 60.8 | 2026-04 |
| gemma 2 27B it | Google DeepMind | 60.2 | 2026-04 |
| Qwen2.5 Coder 32B Instruct | Alibaba / Qwen Team | 60.2 | 2026-04 |
| Qwen2.5 Coder 32B Instruct AWQ | Alibaba / Qwen Team | 60.2 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 60.1 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 58.8 | 2026-04 |
| Codestral (latest) | Mistral AI | 58.5 | 2026-04 |
| Command R+ | Cohere | 58.5 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 58.2 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 58.2 | 2026-04 |
| Qwen2.5 14B Instruct | Alibaba / Qwen Team | 58.2 | 2026-04 |
| Qwen2.5 14B Instruct AWQ | Alibaba / Qwen Team | 58.2 | 2026-04 |
| Qwen3 8B | Alibaba / Qwen Team | 57.8 | 2026-04 |
| Qwen3 8B AWQ | Alibaba / Qwen Team | 57.8 | 2026-04 |
| Qwen3 8B Base | Alibaba / Qwen Team | 57.8 | 2026-04 |
| GPT-4 | OpenAI | 56.2 | 2026-04 |
| Claude Haiku 3 | Anthropic | 55.8 | 2026-04 |
| Yi 1.5 34B | 01.AI | 52.8 | 2026-04 |
| Yi 1.5 34B Chat | 01.AI | 52.8 | 2026-04 |
| Mixtral 8x22B | Mistral AI | 52.3 | 2026-04 |
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 52.3 | 2026-04 |
| Meta Llama 3 70B Instruct | Meta | 52.1 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 52.1 | 2026-04 |
| Command R | Cohere | 50.2 | 2026-04 |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | 50.2 | 2026-04 |
| Qwen2.5 7B | Alibaba / Qwen Team | 50.1 | 2026-04 |
| Qwen2.5 7B Instruct | Alibaba / Qwen Team | 50.1 | 2026-04 |
| Qwen2.5 7B Instruct AWQ | Alibaba / Qwen Team | 50.1 | 2026-04 |
| DeepSeek R1 Distill Llama 8B | DeepSeek | 48.5 | 2026-04 |
| gemma 2 9B | Google DeepMind | 48.5 | 2026-04 |
| gemma 2 9B it | Google DeepMind | 48.5 | 2026-04 |
| Qwen3 4B | Alibaba / Qwen Team | 48.5 | 2026-04 |
| Qwen3 4B Base | Alibaba / Qwen Team | 48.5 | 2026-04 |
| Qwen3 4B Instruct 2507 | Alibaba / Qwen Team | 48.5 | 2026-04 |
| Qwen3 4B Instruct 2507 FP8 | Alibaba / Qwen Team | 48.5 | 2026-04 |
| Mistral Nemo | Mistral AI | 47.8 | 2026-04 |
| Mistral Nemo Base 2407 | Mistral AI | 47.8 | 2026-04 |