The same 250 GSM8K grade-school math problems, human-translated into ten languages, to test whether chain-of-thought reasoning holds up outside English.
unassessed
| Category | math |
|---|---|
| Subcategory | multilingual math reasoning |
| Page status | active |
| Metric | accuracy (exact match on final numeric answer) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 250 |
| Dataset licence | CC-BY-4.0 |
| Publisher | Google Research |
MGSM gives a model the same 250 grade-school arithmetic word problems used in GSM8K, each professionally translated by human annotators into ten languages (Spanish, French, German, Russian, Chinese, Japanese, Thai, Swahili, Bengali, Telugu), alongside the English originals. The model reads one problem in one language and produces a final numeric answer, typically after a worked chain-of-thought solution. The paper frames this as a test of multilingual chain-of-thought reasoning, not of translation quality: the translation is fixed and given to the model, and what is scored is whether multi-step arithmetic reasoning still works once the problem is not in English.
Free-response grade-school word problem in one of eleven languages; the model outputs a final answer as an Arabic numeral, usually after a chain-of-thought solution.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Cerebras-Llama-4-Maverick-17B-128E-Instruct | Meta | 92.3 | 2026-04 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 92.3 | 2026-04 |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | Meta | 92.3 | 2026-04 |
| o3-mini | OpenAI | 92.0 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 91.6 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 91.6 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 91.1 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 91.1 | 2026-04 |
| Meta Llama 3 70B Instruct | Meta | 91.1 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 91.1 | 2026-04 |
| o1-preview | OpenAI | 90.8 | 2026-04 |
| Claude Opus 3 | Anthropic | 90.7 | 2026-04 |
| Llama 4 Scout 17B 16E | Meta | 90.6 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 90.6 | 2026-04 |
| Llama-4-Scout-17B-16E-Instruct-FP8 | Meta | 90.6 | 2026-04 |
| GPT-4o | OpenAI | 90.5 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 90.5 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 90.5 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 90.5 | 2026-04 |
| o1 | OpenAI | 89.3 | 2026-04 |
| o1-mini | OpenAI | 89.3 | 2026-04 |
| o1-pro | OpenAI | 89.3 | 2026-04 |
| GPT-4 Turbo | OpenAI | 88.5 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 87.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 87.5 | 2026-04 |
| GPT-4o mini | OpenAI | 87.0 | 2026-04 |
| Llama 3.2 90B Vision Instruct | Meta | 86.9 | 2026-04 |
| Claude Haiku 3.5 | Anthropic | 85.6 | 2026-04 |
| Claude Haiku 3.5 (latest) | Anthropic | 85.6 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 83.5 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 83.5 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 82.6 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 82.6 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 82.6 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 82.6 | 2026-04 |
| phi 4 | Microsoft | 80.6 | 2026-04 |
| Claude Haiku 3 | Anthropic | 75.1 | 2026-04 |
| GPT-4 | OpenAI | 74.5 | 2026-04 |
| Llama 3.2 11B Vision | Meta | 68.9 | 2026-04 |
| Gemma 3n 4B | Google DeepMind | 67.0 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 63.9 | 2026-04 |
| Llama 3.2 3B | Meta | 58.2 | 2026-04 |
| Llama 3.2 3B Instruct | Meta | 58.2 | 2026-04 |
| GPT-3.5-turbo | OpenAI | 56.3 | 2026-04 |
| Phi 3.5 mini instruct | Microsoft | 47.9 | 2026-04 |