MMLU subject subset: Bar-exam-style fact patterns testing US law, and by far the largest of MMLU's 57 subjects.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | law |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1534 |
| Dataset licence | MIT |
| Publisher | UC Berkeley (original); Center for AI Safety (current host) |
Bar-exam-style fact patterns testing US law across contracts, torts, criminal law, property and evidence, each resolved to a single correct rule application. Questions are four-option multiple-choice, drawn from the MMLU test set's "law" subcategory within the benchmark's "humanities" top-level group, and are graded on the single correct labelled option.
Four-option multiple-choice questions, graded on the single correct labelled option; commonly evaluated 5-shot, consistent with the rest of MMLU.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Opus 4 | Anthropic | 78.2 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 78.2 | 2026-04 |
| GPT-4.1 | OpenAI | 77.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 77.2 | 2026-04 |
| GPT-4o | OpenAI | 76.8 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 76.8 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 76.8 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 76.8 | 2026-04 |
| GPT-4o mini | OpenAI | 76.8 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 75.5 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 75.5 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 75.5 | 2026-04 |
| DeepSeek R1 | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 72.5 | 2026-04 |
| DeepSeek R1 0528 Qwen3 8B | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 Distill Llama 8B | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 1.5B | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | 72.5 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 72.5 | 2026-04 |
| Qwen 3 235B Instruct | Cerebras | 71.8 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 71.8 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 70.5 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 70.5 | 2026-04 |
| Mistral Large 3 | Mistral AI | 70.5 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 70.2 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 70.2 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 70.2 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 70.2 | 2026-04 |
| Gemma 4 26B | Google DeepMind | 69.2 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 68.2 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 68.2 | 2026-04 |
| Llama 3.1 70B | Meta | 67.5 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 67.5 | 2026-04 |
| phi 4 | Microsoft | 65.5 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 65.5 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 65.5 | 2026-04 |
| Meta Llama 3 70B | Meta | 64.1 | 2024-07 |
| Meta Llama 3 70B Instruct | Meta | 64.1 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 64.1 | 2026-04 |
| Nous Hermes 2 Yi 34B | Nous Research | 61.7 | 2024-07 |
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 60.3 | 2026-04 |
| Yi 34B 200K | 01.AI | 59.1 | 2024-07 |
| Yi 1.5 34B 32K | 01.AI | 58.5 | 2024-07 |
| Yi 1.5 34B | 01.AI | 58.0 | 2026-04 |
| Yi 1.5 34B Chat | 01.AI | 57.0 | 2026-04 |
| Yi 1.5 34B Chat 16K | 01.AI | 57.0 | 2024-07 |
| Nous Hermes 2 Mixtral 8x7B DPO | Nous Research | 55.5 | 2024-07 |
| Yi 34B Chat | 01.AI | 55.0 | 2024-07 |
| Mixtral 8x7B Instruct v0.1 | Mistral AI | 54.4 | 2026-04 |
| Mixtral 8x7B v0.1 | Mistral AI | 53.2 | 2026-04 |
| Phi 3 mini 4K instruct | Microsoft | 51.3 | 2024-07 |
| Yi 1.5 9B 32K | 01.AI | 50.7 | 2024-07 |
| Yi 1.5 9B | 01.AI | 50.3 | 2024-07 |
| Nous Hermes 2 SOLAR 10.7B | Nous Research | 50.1 | 2024-07 |
| Yi 1.5 9B Chat 16K | 01.AI | 50.1 | 2024-07 |
| Phi 3 mini 128K instruct | Microsoft | 49.7 | 2024-07 |
| Yi 6B | 01.AI | 49.5 | 2024-07 |
| Yi 6B Chat | 01.AI | 49.5 | 2024-07 |
| Yi 9B | 01.AI | 49.5 | 2024-07 |
| Hermes 2 Theta Llama 3 8B | Nous Research | 48.5 | 2024-07 |
| gemma 7B it | Google DeepMind | 48.1 | 2024-07 |
| Yi 1.5 9B Chat | 01.AI | 48.0 | 2024-07 |
| Meta Llama 3 8B Instruct | Meta | 47.8 | 2024-07 |
| Meta Llama 3 8B Instruct | Nous Research | 47.8 | 2024-07 |
| Yi 1.5 6B Chat | 01.AI | 47.3 | 2024-07 |
| Yi 1.5 6B | 01.AI | 47.1 | 2024-07 |
| Hermes 2 Pro Llama 3 8B | Nous Research | 46.8 | 2024-07 |
| Meta Llama 3 8B | Meta | 46.7 | 2024-07 |
| Meta Llama 3 8B | Nous Research | 46.7 | 2024-07 |
| Mistral 7B v0.3 | Mistral AI | 46.2 | 2024-07 |
| mistral 7B v0.3 bnb 4bit | Unsloth | 46.2 | 2024-07 |
| phi 2 | Microsoft | 43.7 | 2024-07 |
| Mistral 7B Instruct v0.2 | Mistral AI | 43.5 | 2024-07 |
| falcon 40B | TII | 43.2 | 2024-07 |
| Qwen2 1.5B Instruct | Alibaba / Qwen Team | 42.8 | 2024-07 |
| deepseek llm 7B base | DeepSeek | 40.7 | 2024-07 |
| deepseek llm 7B chat | DeepSeek | 40.7 | 2024-07 |
| Qwen2 0.5B Instruct | Alibaba / Qwen Team | 36.2 | 2024-07 |
| chatglm2 6B | Zhipu AI | 35.6 | 2024-07 |
| gemma 2B | Google DeepMind | 34.7 | 2024-07 |
| gemma 2B it | Google DeepMind | 31.9 | 2024-07 |
| deepseek coder 1.3B base | DeepSeek | 29.0 | 2024-07 |
| deepseek coder 1.3B instruct | DeepSeek | 29.0 | 2024-07 |
| deepseek coder 6.7B base | DeepSeek | 28.8 | 2024-07 |
| deepseek coder 6.7B instruct | DeepSeek | 28.4 | 2024-07 |
| OLMo 1B hf | Allen AI | 24.1 | 2024-07 |