8.5K grade-school math word problems scored by exact match on the final numeric answer, testing multi-step arithmetic reasoning.
unassessed
| Category | math |
|---|---|
| Subcategory | grade-school arithmetic word problems |
| Page status | active |
| Metric | accuracy (exact-match final answer) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 8792 |
| Dataset licence | MIT |
| Publisher | OpenAI |
GSM8K tests whether a model can solve short, English-language math word problems that take two to eight linked arithmetic steps to reach an answer, such as working out a total after several purchases and discounts. OpenAI wrote the problems so that a bright middle-school student could solve every one of them by hand, deliberately excluding algebra, calculus or other advanced mathematics. The skill under test is multi-step arithmetic planning and consistent calculation rather than mathematical sophistication.
Free-response word problem in English; the model produces a reasoning chain ending in one final numeric answer, extracted after a `####` marker and graded by exact match.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Kimi K2 0711 | Moonshot AI | 97.3 | 2026-04 |
| Kimi K2 0905 | Moonshot AI | 97.3 | 2026-04 |
| Kimi K2 Thinking | Moonshot AI | 97.3 | 2026-04 |
| Kimi K2 Thinking Turbo | Moonshot AI | 97.3 | 2026-04 |
| Kimi K2 Turbo | Moonshot AI | 97.3 | 2026-04 |
| Kimi K2.5 | Moonshot AI | 97.3 | 2026-04 |
| Kimi K2.5 NVFP4 | NVIDIA | 97.3 | 2026-04 |
| o1 | OpenAI | 97.1 | 2026-04 |
| o1-mini | OpenAI | 97.1 | 2026-04 |
| o1-pro | OpenAI | 97.1 | 2026-04 |
| Llama 3.1 405B | Meta | 96.8 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 96.8 | 2026-04 |
| Llama 3.1 405B Instruct FP8 | Meta | 96.8 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 96.4 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 96.4 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 95.9 | 2026-04 |
| Claude Opus 3 | Anthropic | 95.0 | 2026-04 |
| GPT-4o | OpenAI | 95.0 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 95.0 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 95.0 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 95.0 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 95.0 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 95.0 | 2026-04 |
| Meta Llama 3 70B Instruct | Meta | 95.0 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 95.0 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 94.4 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 94.4 | 2026-04 |
| phi 4 | Microsoft | 93.5 | 2026-04 |
| GPT-4 Turbo | OpenAI | 93.0 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 93.0 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 92.3 | 2026-04 |
| GPT-4 | OpenAI | 92.0 | 2026-04 |
| GPT-4o mini | OpenAI | 92.0 | 2026-04 |
| Qwen2.5 7B | Alibaba / Qwen Team | 91.6 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 90.8 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 90.8 | 2026-04 |
| DeepSeek Chat | DeepSeek | 89.3 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 89.3 | 2026-04 |
| DeepSeek V3 | DeepSeek | 89.3 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 89.3 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 89.3 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 89.3 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 89.3 | 2026-04 |
| Gemma 3 4B | Google DeepMind | 89.2 | 2026-04 |
| gemma 3 4B pt | Google DeepMind | 89.2 | 2026-04 |
| Claude Haiku 3 | Anthropic | 88.9 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 88.6 | 2026-04 |
| Mixtral 8x22B | Mistral AI | 88.0 | 2026-04 |
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 88.0 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 86.2 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 86.2 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 86.2 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 86.2 | 2026-04 |
| Phi 3.5 mini instruct | Microsoft | 86.2 | 2026-04 |
| Meta Llama 3 70B | Meta | 85.4 | 2024-07 |
| granite 3.3 8B instruct | IBM | 80.9 | 2026-04 |
| GPT-3.5-turbo | OpenAI | 80.0 | 2026-04 |
| Llama 3.2 3B | Meta | 77.7 | 2026-04 |
| Llama 3.2 3B Instruct | Meta | 77.7 | 2026-04 |
| Phi 3 mini 4K instruct | Microsoft | 74.5 | 2024-07 |
| gemma 2 27B it | Google DeepMind | 74.0 | 2026-04 |
| Yi 1.5 34B | 01.AI | 73.2 | 2026-04 |
| Yi 1.5 9B Chat | 01.AI | 71.9 | 2024-07 |
| Nous Hermes 2 Mixtral 8x7B DPO | Nous Research | 71.6 | 2024-07 |
| Yi 1.5 34B Chat | 01.AI | 71.6 | 2026-04 |
| Yi 1.5 34B Chat 16K | 01.AI | 71.6 | 2024-07 |
| Command R+ | Cohere | 70.7 | 2026-04 |
| Hermes 2 Theta Llama 3 8B | Nous Research | 70.4 | 2024-07 |
| Nous Hermes 2 Yi 34B | Nous Research | 70.1 | 2024-07 |
| Phi 3 mini 128K instruct | Microsoft | 69.5 | 2024-07 |
| Nous Hermes 2 SOLAR 10.7B | Nous Research | 69.4 | 2024-07 |
| Meta Llama 3 8B Instruct | Meta | 68.7 | 2024-07 |
| Meta Llama 3 8B Instruct | Nous Research | 68.7 | 2024-07 |
| gemma 2 9B | Google DeepMind | 68.6 | 2026-04 |
| gemma 2 9B it | Google DeepMind | 68.6 | 2026-04 |
| Hermes 2 Pro Llama 3 8B | Nous Research | 67.9 | 2024-07 |
| Yi 1.5 6B Chat | 01.AI | 67.1 | 2024-07 |
| gemma 3 1B it | Google DeepMind | 62.8 | 2026-04 |
| gemma 3 1B pt | Google DeepMind | 62.8 | 2026-04 |
| Yi 1.5 9B | 01.AI | 62.7 | 2024-07 |
| Mixtral 8x7B Instruct v0.1 | Mistral AI | 61.1 | 2026-04 |
| Yi 1.5 9B Chat 16K | 01.AI | 59.7 | 2024-07 |
| Mixtral 8x7B v0.1 | Mistral AI | 57.6 | 2026-04 |
| Qwen2 1.5B Instruct | Alibaba / Qwen Team | 55.8 | 2024-07 |
| phi 2 | Microsoft | 55.0 | 2024-07 |
| Llama 2 70B hf | Meta | 54.1 | 2024-07 |
| gemma 7B it | Google DeepMind | 52.8 | 2024-07 |
| Yi 34B | 01.AI | 50.6 | 2024-07 |
| Yi 1.5 6B | 01.AI | 49.8 | 2024-07 |
| Yi 9B | 01.AI | 49.0 | 2024-07 |
| deepseek llm 7B base | DeepSeek | 45.9 | 2024-07 |
| deepseek llm 7B chat | DeepSeek | 45.9 | 2024-07 |
| Meta Llama 3 8B | Meta | 45.2 | 2024-07 |
| Meta Llama 3 8B | Nous Research | 45.2 | 2024-07 |
| Mistral 7B Instruct v0.2 | Mistral AI | 40.0 | 2024-07 |
| Mistral 7B v0.1 | Mistral AI | 37.1 | 2024-07 |
| Qwen2 0.5B Instruct | Alibaba / Qwen Team | 35.7 | 2024-07 |
| Yi 34B 200K | 01.AI | 34.9 | 2024-07 |
| Mistral 7B v0.3 | Mistral AI | 34.5 | 2024-07 |
| mistral 7B v0.3 bnb 4bit | Unsloth | 34.5 | 2024-07 |
| falcon 40B instruct | TII | 34.3 | 2024-07 |
| Yi 34B Chat | 01.AI | 31.9 | 2024-07 |
| Yi 6B 200K | 01.AI | 30.3 | 2024-07 |
| deepseek coder 6.7B instruct | DeepSeek | 26.8 | 2024-07 |
| Llama 2 70B chat hf | Meta | 26.7 | 2024-07 |
| Llama 2 13B hf | Meta | 22.8 | 2024-07 |
| Llama 2 13B hf | Nous Research | 22.8 | 2024-07 |
| CodeLlama 34B Instruct hf | Meta | 21.6 | 2026-04 |
| deepseek coder 6.7B base | DeepSeek | 18.0 | 2024-07 |
| gemma 2B | Google DeepMind | 16.9 | 2024-07 |
| Llama 2 13B chat hf | Meta | 15.2 | 2024-07 |
| Llama 2 7B hf | Meta | 14.5 | 2024-07 |
| Llama 2 7B hf | Nous Research | 14.5 | 2024-07 |
| Mistral 7B Instruct v0.1 | Mistral AI | 14.3 | 2024-07 |
| Yi 6B | 01.AI | 12.1 | 2024-07 |
| Yi 6B Chat | 01.AI | 12.1 | 2024-07 |
| Llama 2 7B chat hf | Meta | 7.4 | 2024-07 |
| Llama 2 7B chat hf | Nous Research | 7.4 | 2024-07 |
| Nous Hermes llama 2 7B | Nous Research | 5.8 | 2024-07 |
| CodeLlama 7B hf | Meta | 5.5 | 2024-07 |
| CodeLlama 7B Instruct hf | Meta | 5.5 | 2024-07 |
| gemma 2B it | Google DeepMind | 5.5 | 2024-07 |
| falcon 7B instruct | TII | 4.7 | 2024-07 |
| falcon 7B | TII | 4.6 | 2024-07 |
| Baichuan 7B | Baichuan | 2.8 | 2024-07 |
| OLMo 1B hf | Allen AI | 1.9 | 2024-07 |
| deepseek coder 1.3B base | DeepSeek | 1.1 | 2024-07 |
| deepseek coder 1.3B instruct | DeepSeek | 1.1 | 2024-07 |
| Yi 1.5 34B 32K | 01.AI | 0.0 | 2024-07 |
| Yi 1.5 9B 32K | 01.AI | 0.0 | 2024-07 |