The computer science subcategory of MMLU: a rollup of four subjects, used by publishers that report MMLU at a coarser grain than all 57 subjects.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | computer science |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 412 |
| Dataset licence | MIT |
| Publisher | UC Berkeley (original); Center for AI Safety (current host) |
This id does not correspond to a single dataset config in the Hugging Face mirror of MMLU. It corresponds to the "computer science" subcategory the benchmark's authors define in the repository's categories.py, which pools four subjects: College Computer Science (algorithms and computability at undergraduate level), High School Computer Science (introductory programming concepts), Computer Security (cryptography and systems security) and Machine Learning. Some publishers report MMLU broken down by this kind of subcategory rather than by all 57 individual subjects; this id captures scores reported at that grain.
Four-option multiple-choice questions pooled from four underlying MMLU subjects, graded on the single correct labelled option. How a given publisher averages the four subjects into one number -- an unweighted mean of per-subject accuracy, or a single accuracy over the pooled question set -- is not documented and not established here.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Opus 4 | Anthropic | 88.8 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 88.8 | 2026-04 |
| GPT-4.1 | OpenAI | 88.2 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 87.8 | 2026-04 |
| GPT-4o | OpenAI | 87.5 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 87.5 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 87.5 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 87.5 | 2026-04 |
| GPT-4o mini | OpenAI | 87.5 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 86.2 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 86.2 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 86.2 | 2026-04 |
| DeepSeek R1 | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 85.5 | 2026-04 |
| DeepSeek R1 0528 Qwen3 8B | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 Distill Llama 8B | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 1.5B | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 85.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | 85.5 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 85.5 | 2026-04 |
| Qwen 3 235B Instruct | Cerebras | 84.5 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 84.5 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 82.5 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 82.5 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 82.5 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 82.5 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 82.2 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 82.2 | 2026-04 |
| Mistral Large 3 | Mistral AI | 82.2 | 2026-04 |
| Gemma 4 26B | Google DeepMind | 81.2 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 80.5 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 80.5 | 2026-04 |
| Llama 3.1 70B | Meta | 79.8 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 79.8 | 2026-04 |
| phi 4 | Microsoft | 78.5 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 78.5 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 78.5 | 2026-04 |