Disputed key on 40 frontier cards: possibly a mis-keyed MMLU-Pro Physics score, possibly a classic-MMLU STEM subcategory rollup. Neither reading is confirmed; pending a card re-key.
unassessed
| Category | knowledge |
|---|---|
| Page status | unknown |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | MIT |
Not established. This key appears on 40 frontier model cards in this repository with no notes on its source. It may be a mis-keyed MMLU-Pro Physics category score (MMLU-Pro has a ten-option Physics category; the original 57-subject MMLU has no subject or dataset config named plain "physics"), or it may be the "physics" STEM subcategory the original MMLU authors define in categories.py, which pools four classic subjects. See "What it measures" below for the evidence on each reading and why neither is confirmed.
Not established -- it depends on which reading is correct. A ten-option, chain-of-thought format if this is an MMLU-Pro category score, or the classic four-option format if this is a categories.py subcategory rollup of four MMLU subjects.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Opus 4 | Anthropic | 85.1 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 85.1 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 84.8 | 2026-04 |
| GPT-4.1 | OpenAI | 84.2 | 2026-04 |
| GPT-4o | OpenAI | 83.5 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 83.5 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 83.5 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 83.5 | 2026-04 |
| GPT-4o mini | OpenAI | 83.5 | 2026-04 |
| DeepSeek R1 | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 82.8 | 2026-04 |
| DeepSeek R1 0528 Qwen3 8B | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 Distill Llama 8B | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 Distill Qwen 1.5B | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 82.8 | 2026-04 |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | 82.8 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 82.8 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 82.5 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 82.5 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 82.5 | 2026-04 |
| Qwen 3 235B Instruct | Cerebras | 81.5 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 81.5 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 78.8 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 78.8 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 78.8 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 78.8 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 78.2 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 78.2 | 2026-04 |
| Mistral Large 3 | Mistral AI | 78.2 | 2026-04 |
| Gemma 4 26B | Google DeepMind | 77.5 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 76.8 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 76.8 | 2026-04 |
| Llama 3.1 70B | Meta | 75.5 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 75.5 | 2026-04 |
| phi 4 | Microsoft | 74.2 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 74.2 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 74.2 | 2026-04 |