A bilingual, 3,709-problem suite spanning five education stages from arithmetic to college, each scored separately on theory recall and applied problem-solving using circular multiple-choice evaluation.
unassessed
| Category | math |
|---|---|
| Subcategory | hierarchical, bilingual theory-and-application mathematics evaluation |
| Page status | active |
| Metric | accuracy under Circular Evaluation (CE) for multiple-choice splits; plain accuracy for cloze splits |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 3709 |
| Dataset licence | Apache-2.0 |
| Publisher | Shanghai AI Laboratory, with contributing authors at Beihang University and Nanjing University |
MathBench tests mathematics proficiency at five difficulty stages that mirror school progression: arithmetic, primary, middle, high school and college. Each stage (other than arithmetic) is tested twice, on two different things: "Application" problems (can the model solve a problem at that level) and "Theory" questions (does the model know the underlying concepts and definitions at that level, independent of solving anything). The authors built this specifically because prior math benchmarks like GSM8K, in their view, gave only a single undifferentiated difficulty signal. Arithmetic and primary-level application problems are free-response word problems; every other stage and the theory questions throughout are four-option multiple choice. All stages except arithmetic are presented in both Chinese and English.
Two formats depending on stage and split: free-response cloze problems (arithmetic, primary application) graded on the final extracted number, and four-option multiple-choice questions (middle/high/college application, and theory questions at every stage) graded with Circular Evaluation -- the same question is re-asked with its option order rotated across 4 rounds (CE-4), and a model is only scored correct on that question if all 4 rotations are answered correctly.
No model card in ModelSpec reports this benchmark yet.