MGSM (Multilingual Grade School Math)

The same 250 GSM8K grade-school math problems, human-translated into ten languages, to test whether chain-of-thought reasoning holds up outside English.

Also known as: Multilingual Grade School Math Benchmark, Multilingual GSM8K

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategorymultilingual math reasoning
Page statusactive
Metricaccuracy (exact match on final numeric answer)
Directionhigher_is_better
Unit%
Dataset size250
Dataset licenceCC-BY-4.0
PublisherGoogle Research

What it measures

MGSM gives a model the same 250 grade-school arithmetic word problems used in GSM8K, each professionally translated by human annotators into ten languages (Spanish, French, German, Russian, Chinese, Japanese, Thai, Swahili, Bengali, Telugu), alongside the English originals. The model reads one problem in one language and produces a final numeric answer, typically after a worked chain-of-thought solution. The paper frames this as a test of multilingual chain-of-thought reasoning, not of translation quality: the translation is fixed and given to the model, and what is scored is whether multi-step arithmetic reasoning still works once the problem is not in English.

Task format

Free-response grade-school word problem in one of eleven languages; the model outputs a final answer as an Arabic numeral, usually after a chain-of-thought solution.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Cerebras-Llama-4-Maverick-17B-128E-InstructMeta92.32026-04
Llama 4 Maverick 17B 128E InstructMeta92.32026-04
Llama-4-Maverick-17B-128E-Instruct-FP8Meta92.32026-04
o3-miniOpenAI92.02026-04
Claude Sonnet 3.5Anthropic91.62026-04
Claude Sonnet 3.5 v2Anthropic91.62026-04
Llama 3.3 70B Instruct NVFP4NVIDIA91.12026-04
Llama-3.3-70B-InstructMeta91.12026-04
Meta Llama 3 70B InstructMeta91.12026-04
Meta Llama 3 70B InstructNous Research91.12026-04
o1-previewOpenAI90.82026-04
Claude Opus 3Anthropic90.72026-04
Llama 4 Scout 17B 16EMeta90.62026-04
Llama 4 Scout 17B 16E InstructMeta90.62026-04
Llama-4-Scout-17B-16E-Instruct-FP8Meta90.62026-04
GPT-4oOpenAI90.52026-04
GPT-4o (2024-05-13)OpenAI90.52026-04
GPT-4o (2024-08-06)OpenAI90.52026-04
GPT-4o (2024-11-20)OpenAI90.52026-04
o1OpenAI89.32026-04
o1-miniOpenAI89.32026-04
o1-proOpenAI89.32026-04
GPT-4 TurboOpenAI88.52026-04
Gemini 1.5 ProGoogle DeepMind87.52026-04
Gemini 2.5 ProGoogle DeepMind87.52026-04
GPT-4o miniOpenAI87.02026-04
Llama 3.2 90B Vision InstructMeta86.92026-04
Claude Haiku 3.5Anthropic85.62026-04
Claude Haiku 3.5 (latest)Anthropic85.62026-04
Claude Sonnet 3Anthropic83.52026-04
Qwen3 235B-A22BAlibaba / Qwen Team83.52026-04
Gemini 1.5 FlashGoogle DeepMind82.62026-04
Gemini 1.5 Flash-8BGoogle DeepMind82.62026-04
Gemini 2.0 FlashGoogle DeepMind82.62026-04
Gemini 2.5 FlashGoogle DeepMind82.62026-04
phi 4Microsoft80.62026-04
Claude Haiku 3Anthropic75.12026-04
GPT-4OpenAI74.52026-04
Llama 3.2 11B VisionMeta68.92026-04
Gemma 3n 4BGoogle DeepMind67.02026-04
Phi 4 mini instructMicrosoft63.92026-04
Llama 3.2 3BMeta58.22026-04
Llama 3.2 3B InstructMeta58.22026-04
GPT-3.5-turboOpenAI56.32026-04
Phi 3.5 mini instructMicrosoft47.92026-04

Data

This page as JSON · Edit on GitHub