MathBench

A bilingual, 3,709-problem suite spanning five education stages from arithmetic to college, each scored separately on theory recall and applied problem-solving using circular multiple-choice evaluation.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategoryhierarchical, bilingual theory-and-application mathematics evaluation
Page statusactive
Metricaccuracy under Circular Evaluation (CE) for multiple-choice splits; plain accuracy for cloze splits
Directionhigher_is_better
Unit%
Dataset size3709
Dataset licenceApache-2.0
PublisherShanghai AI Laboratory, with contributing authors at Beihang University and Nanjing University

What it measures

MathBench tests mathematics proficiency at five difficulty stages that mirror school progression: arithmetic, primary, middle, high school and college. Each stage (other than arithmetic) is tested twice, on two different things: "Application" problems (can the model solve a problem at that level) and "Theory" questions (does the model know the underlying concepts and definitions at that level, independent of solving anything). The authors built this specifically because prior math benchmarks like GSM8K, in their view, gave only a single undifferentiated difficulty signal. Arithmetic and primary-level application problems are free-response word problems; every other stage and the theory questions throughout are four-option multiple choice. All stages except arithmetic are presented in both Chinese and English.

Task format

Two formats depending on stage and split: free-response cloze problems (arithmetic, primary application) graded on the final extracted number, and four-option multiple-choice questions (middle/high/college application, and theory questions at every stage) graded with Circular Evaluation -- the same question is re-asked with its option order rotated across 4 rounds (CE-4), and a model is only scored correct on that question if all 4 rotations are answered correctly.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub