MMLU: Physics (unresolved key)

Disputed key on 40 frontier cards: possibly a mis-keyed MMLU-Pro Physics score, possibly a classic-MMLU STEM subcategory rollup. Neither reading is confirmed; pending a card re-key.

Also known as: mmlu_pro_physics

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Page statusunknown
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset licenceMIT

What it measures

Not established. This key appears on 40 frontier model cards in this repository with no notes on its source. It may be a mis-keyed MMLU-Pro Physics category score (MMLU-Pro has a ten-option Physics category; the original 57-subject MMLU has no subject or dataset config named plain "physics"), or it may be the "physics" STEM subcategory the original MMLU authors define in categories.py, which pools four classic subjects. See "What it measures" below for the evidence on each reading and why neither is confirmed.

Task format

Not established -- it depends on which reading is correct. A ten-option, chain-of-thought format if this is an MMLU-Pro category score, or the classic four-option format if this is a categories.py subcategory rollup of four MMLU subjects.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4Anthropic85.12026-04
Claude Opus 4.6Anthropic85.12026-04
Gemini 2.5 ProGoogle DeepMind84.82026-04
GPT-4.1OpenAI84.22026-04
GPT-4oOpenAI83.52026-04
GPT-4o (2024-05-13)OpenAI83.52026-04
GPT-4o (2024-08-06)OpenAI83.52026-04
GPT-4o (2024-11-20)OpenAI83.52026-04
GPT-4o miniOpenAI83.52026-04
DeepSeek R1DeepSeek82.82026-04
DeepSeek R1 0528DeepSeek82.82026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA82.82026-04
DeepSeek R1 0528 Qwen3 8BDeepSeek82.82026-04
DeepSeek R1 Distill Llama 70BDeepSeek82.82026-04
DeepSeek R1 Distill Llama 8BDeepSeek82.82026-04
DeepSeek R1 Distill Qwen 1.5BDeepSeek82.82026-04
DeepSeek R1 Distill Qwen 14BDeepSeek82.82026-04
DeepSeek R1 Distill Qwen 32BDeepSeek82.82026-04
DeepSeek R1 Distill Qwen 7BDeepSeek82.82026-04
DeepSeek ReasonerDeepSeek82.82026-04
Claude Sonnet 4Anthropic82.52026-04
Claude Sonnet 4.5Anthropic82.52026-04
Claude Sonnet 4.5 (latest)Anthropic82.52026-04
Qwen 3 235B InstructCerebras81.52026-04
Qwen3 235B-A22BAlibaba / Qwen Team81.52026-04
Gemma 4 31BGoogle DeepMind78.82026-04
gemma 4 31B itGoogle DeepMind78.82026-04
gemma 4 31B it GGUFUnsloth78.82026-04
Gemma 4 31B IT NVFP4NVIDIA78.82026-04
Mistral Large (latest)Mistral AI78.22026-04
Mistral Large 2.1Mistral AI78.22026-04
Mistral Large 3Mistral AI78.22026-04
Gemma 4 26BGoogle DeepMind77.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA76.82026-04
Llama-3.3-70B-InstructMeta76.82026-04
Llama 3.1 70BMeta75.52026-04
Llama 3.1 70B InstructMeta75.52026-04
phi 4Microsoft74.22026-04
Phi 4 mini instructMicrosoft74.22026-04
Phi 4 multimodal instructMicrosoft74.22026-04

Data

This page as JSON · Edit on GitHub