MMLU-Pro

A harder, ten-option successor to MMLU with about 12,000 reasoning-heavy questions across 14 categories, built to restore headroom lost to MMLU's saturation at the frontier.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorymultitask academic and professional knowledge (reasoning-augmented)
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size12032
Dataset licenceMIT
PublisherTIGER Lab, University of Waterloo (with co-authors at the University of Toronto and Carnegie Mellon University)

What it measures

MMLU-Pro gives a model a question and up to ten labelled answer options (most questions use all ten; a small number carry fewer after manual review removed unreasonable distractors), drawn from 14 broad categories spanning STEM, humanities, social sciences and business/health/other topics: Biology, Business, Chemistry, Computer Science, Economics, Engineering, Health, History, Law, Math, Philosophy, Physics, Psychology and Other. The model must select the single correct option. Roughly 57% of the questions are a difficulty-filtered subset of the original MMLU test set (with "trivial and ambiguous" items removed); the remaining 43% are newly written from a STEM question website, TheoremQA and SciBench, with GPT-4-generated distractors expanding every question from four options toward ten, reviewed afterward by a panel of over ten subject-matter experts.

Task format

Multiple-choice question answering with up to ten labelled options, graded on the single option the model selects. The reference protocol is 5-shot with chain-of-thought prompting, with fewshot examples drawn from a dedicated 70-question validation split; scoring extracts the answer letter from free-form generated text (for example via a regex matching "the answer is (X)") rather than comparing option log-likelihoods, because the authors found direct/log-likelihood scoring under-performs chain-of-thought by up to 19 points on this dataset -- the opposite of the original MMLU. Across 24 prompt styles the authors tested, score sensitivity to prompt wording fell from 4-5% on MMLU to about 2% on MMLU-Pro.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
GPT-5.1OpenAI86.02026-04
GPT-5.1 ChatOpenAI86.02026-04
GPT-5.1 CodexOpenAI86.02026-04
GPT-5.1 Codex MaxOpenAI86.02026-04
GPT-5.1 Codex miniOpenAI86.02026-04
o3-proOpenAI85.52026-04
Claude Mythos PreviewAnthropic85.22026-04
Claude Opus 4Anthropic85.22026-04
Claude Opus 4.6Anthropic85.22026-04
GPT-5OpenAI84.82026-04
GPT-5 Chat (latest)OpenAI84.82026-04
GPT-5 ProOpenAI84.82026-04
GPT-5-CodexOpenAI84.82026-04
GPT-5.2OpenAI84.82026-04
GPT-5.2 ChatOpenAI84.82026-04
GPT-5.2 CodexOpenAI84.82026-04
GPT-5.2 ProOpenAI84.82026-04
GPT-5.3 Chat (latest)OpenAI84.82026-04
GPT-5.3 CodexOpenAI84.82026-04
GPT-5.3 Codex SparkOpenAI84.82026-04
GPT-5.4OpenAI84.82026-04
GPT-5.4 miniOpenAI84.82026-04
GPT-5.4 nanoOpenAI84.82026-04
GPT-5.4 ProOpenAI84.82026-04
Claude Opus 4.5Anthropic84.52026-04
Claude Opus 4.5 (latest)Anthropic84.52026-04
o3OpenAI84.12026-04
o3-deep-researchOpenAI84.12026-04
Gemini 2.5 ProGoogle DeepMind84.02026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind84.02026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind84.02026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind84.02026-04
Claude Opus 4.1Anthropic83.82026-04
Claude Opus 4.1 (latest)Anthropic83.82026-04
Claude Sonnet 4Anthropic83.02026-04
Claude Sonnet 4.6Anthropic83.02026-04
Grok 4xAI82.52026-04
Grok 4 FastxAI82.52026-04
Grok 4 Fast (Non-Reasoning)xAI82.52026-04
Grok 4.1 FastxAI82.52026-04
Grok 4.1 Fast (Non-Reasoning)xAI82.52026-04
Grok 4.20 (Non-Reasoning)xAI82.52026-04
Grok 4.20 (Reasoning)xAI82.52026-04
Grok 4.20 Multi-AgentxAI82.52026-04
Claude Opus 4 (latest)Anthropic82.12026-04
o1-proOpenAI82.02026-04
Claude Sonnet 4.5Anthropic81.72026-04
Claude Sonnet 4.5 (latest)Anthropic81.72026-04
DeepSeek R1 0528DeepSeek81.02026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA81.02026-04
o4-miniOpenAI81.02026-04
o4-mini-deep-researchOpenAI81.02026-04
Qwen3 235B-A22BAlibaba / Qwen Team80.22026-04
GPT-4.1OpenAI80.12026-04
DeepSeek R1DeepSeek79.82026-04
DeepSeek ReasonerDeepSeek79.82026-04
GPT-5 MiniOpenAI79.52026-04
o3-miniOpenAI79.02026-04
DeepSeek V3.2DeepSeek78.82026-04
DeepSeek V3.2 ExpDeepSeek78.82026-04
Claude Sonnet 4 (latest)Anthropic78.52026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team78.52026-04
Gemini 2.5 FlashGoogle DeepMind78.22026-04
Gemini 2.5 Flash ImageGoogle DeepMind78.22026-04
Gemini 2.5 Flash Image (Preview)Google DeepMind78.22026-04
Gemini 2.5 Flash LiteGoogle DeepMind78.22026-04
Gemini 2.5 Flash Lite Preview 06-17Google DeepMind78.22026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind78.22026-04
Gemini 2.5 Flash Preview 04-17Google DeepMind78.22026-04
Gemini 2.5 Flash Preview 05-20Google DeepMind78.22026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind78.22026-04
Gemini 2.5 Flash Preview TTSGoogle DeepMind78.22026-04
Claude Sonnet 3.7Anthropic78.02026-04
o1OpenAI78.02026-04
o1-previewOpenAI78.02026-04
DeepSeek V3.1DeepSeek77.22026-04
Claude Sonnet 3.5Anthropic76.22026-04
Claude Sonnet 3.5 v2Anthropic76.22026-04
Grok 3xAI76.02026-04
Grok 3 FastxAI76.02026-04
Grok 3 Fast LatestxAI76.02026-04
Grok 3 LatestxAI76.02026-04
DeepSeek ChatDeepSeek75.52026-04
DeepSeek V2DeepSeek75.52026-04
DeepSeek V2 LiteDeepSeek75.52026-04
DeepSeek V2 Lite ChatDeepSeek75.52026-04
DeepSeek V3DeepSeek75.52026-04
DeepSeek V3 0324DeepSeek75.52026-04
Claude Haiku 4.5Anthropic75.22026-04
Claude Haiku 4.5 (latest)Anthropic75.22026-04
Qwen3 32BAlibaba / Qwen Team74.82026-04
Qwen3 32B AWQAlibaba / Qwen Team74.82026-04
Qwen3 32B NVFP4NVIDIA74.82026-04
GPT-4.1 miniOpenAI74.22026-04
Magistral Medium (latest)Mistral AI74.22026-04
Gemma 4 31BGoogle DeepMind74.12026-04
gemma 4 31B itGoogle DeepMind74.12026-04
gemma 4 31B it GGUFUnsloth74.12026-04
Gemma 4 31B IT NVFP4NVIDIA74.12026-04
Gemini 2.0 FlashGoogle DeepMind73.52026-04
Llama 4 Maverick 17B 128E InstructMeta73.52026-04
Llama-4-Maverick-17B-128E-Instruct-FP8Meta73.52026-04
GPT-4oOpenAI72.62026-04
GPT-4o (2024-05-13)OpenAI72.62026-04
GPT-4o (2024-08-06)OpenAI72.62026-04
GPT-4o (2024-11-20)OpenAI72.62026-04
o1-miniOpenAI72.52026-04
QwQ PlusAlibaba / Qwen Team72.52026-04
Gemma 4 26BGoogle DeepMind72.32026-04
gemma 4 26B A4B itGoogle DeepMind72.32026-04
gemma 4 26B A4B it GGUFUnsloth72.32026-04
Gemini 1.5 ProGoogle DeepMind72.02026-04
Qwen2.5 72B InstructAlibaba / Qwen Team71.52026-04
GPT-5 NanoOpenAI70.32026-04
Grok 3 MinixAI70.22026-04
Grok 3 Mini FastxAI70.22026-04
Grok 3 Mini Fast LatestxAI70.22026-04
Grok 3 Mini LatestxAI70.22026-04
Mistral Large (latest)Mistral AI69.82026-04
Mistral Large 2.1Mistral AI69.82026-04
Mistral Large 3Mistral AI69.82026-04
Claude Opus 3Anthropic68.52026-04
phi 4Microsoft68.52026-04
Phi 4 multimodal instructMicrosoft68.52026-04
Qwen3 30B A3B Instruct 2507Alibaba / Qwen Team68.52026-04
Qwen3 30B A3B NVFP4NVIDIA68.52026-04
Qwen3 30B-A3BAlibaba / Qwen Team68.52026-04
DeepSeek R1 Distill Llama 70BDeepSeek68.32026-04
Llama 4 Scout 17B 16EMeta68.22026-04
Llama 4 Scout 17B 16E InstructMeta68.22026-04
Llama-4-Scout-17B-16E-Instruct-FP8Meta68.22026-04
Gemma 3 27BGoogle DeepMind67.52026-04
Llama 3.1 405BMeta67.52026-04
Llama 3.1 405B FP8Meta67.52026-04
Llama 3.1 405B InstructMeta67.52026-04
Llama 3.1 405B Instruct FP8Meta67.52026-04
GPT-4 TurboOpenAI67.42026-04
GPT-4o miniOpenAI66.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA66.52026-04
Llama-3.3-70B-InstructMeta66.52026-04
DeepSeek R1 Distill Qwen 32BDeepSeek65.52026-04
Grok 2xAI65.52026-04
Grok 2 (1212)xAI65.52026-04
Grok 2 LatestxAI65.52026-04
Grok 2 VisionxAI65.52026-04
Grok 2 Vision (1212)xAI65.52026-04
Grok 2 Vision LatestxAI65.52026-04
Claude Haiku 3.5Anthropic65.32026-04
Claude Haiku 3.5 (latest)Anthropic65.32026-04
Qwen2.5 32B InstructAlibaba / Qwen Team65.22026-04
Qwen2.5 32B Instruct AWQAlibaba / Qwen Team65.22026-04
Gemini 2.0 Flash LiteGoogle DeepMind64.82026-04
Magistral SmallMistral AI64.52026-04
Magistral Small 2506Mistral AI64.52026-04
Qwen3 14BAlibaba / Qwen Team64.22026-04
Qwen3 14B AWQAlibaba / Qwen Team64.22026-04
Qwen3 14B NVFP4NVIDIA64.22026-04
GPT-4.1 nanoOpenAI64.12026-04
Command ACohere63.82026-04
Command A ReasoningCohere63.82026-04
Gemini 1.5 FlashGoogle DeepMind63.52026-04
Gemini 1.5 Flash-8BGoogle DeepMind63.52026-04
Llama 3.1 70BMeta60.82026-04
Llama 3.1 70B InstructMeta60.82026-04
gemma 2 27B itGoogle DeepMind60.22026-04
Qwen2.5 Coder 32B InstructAlibaba / Qwen Team60.22026-04
Qwen2.5 Coder 32B Instruct AWQAlibaba / Qwen Team60.22026-04
Claude Sonnet 3Anthropic60.12026-04
DeepSeek R1 Distill Qwen 14BDeepSeek58.82026-04
Codestral (latest)Mistral AI58.52026-04
Command R+Cohere58.52026-04
Gemma 3 12BGoogle DeepMind58.22026-04
Phi 4 mini instructMicrosoft58.22026-04
Qwen2.5 14B InstructAlibaba / Qwen Team58.22026-04
Qwen2.5 14B Instruct AWQAlibaba / Qwen Team58.22026-04
Qwen3 8BAlibaba / Qwen Team57.82026-04
Qwen3 8B AWQAlibaba / Qwen Team57.82026-04
Qwen3 8B BaseAlibaba / Qwen Team57.82026-04
GPT-4OpenAI56.22026-04
Claude Haiku 3Anthropic55.82026-04
Yi 1.5 34B01.AI52.82026-04
Yi 1.5 34B Chat01.AI52.82026-04
Mixtral 8x22BMistral AI52.32026-04
Mixtral 8x22B Instruct v0.1Mistral AI52.32026-04
Meta Llama 3 70B InstructMeta52.12026-04
Meta Llama 3 70B InstructNous Research52.12026-04
Command RCohere50.22026-04
DeepSeek R1 Distill Qwen 7BDeepSeek50.22026-04
Qwen2.5 7BAlibaba / Qwen Team50.12026-04
Qwen2.5 7B InstructAlibaba / Qwen Team50.12026-04
Qwen2.5 7B Instruct AWQAlibaba / Qwen Team50.12026-04
DeepSeek R1 Distill Llama 8BDeepSeek48.52026-04
gemma 2 9BGoogle DeepMind48.52026-04
gemma 2 9B itGoogle DeepMind48.52026-04
Qwen3 4BAlibaba / Qwen Team48.52026-04
Qwen3 4B BaseAlibaba / Qwen Team48.52026-04
Qwen3 4B Instruct 2507Alibaba / Qwen Team48.52026-04
Qwen3 4B Instruct 2507 FP8Alibaba / Qwen Team48.52026-04
Mistral NemoMistral AI47.82026-04
Mistral Nemo Base 2407Mistral AI47.82026-04

Showing the top 200 of 354.

Data

This page as JSON · Edit on GitHub