A 350-problem bilingual Olympiad math suite: 200 numeric easy/hard items plus 150 Lean 4 proofs, after AIME and MATH stopped separating frontier models.
unassessed
| Category | math |
|---|---|
| Subcategory | bilingual olympiad mathematics with numeric answers and Lean proofs |
| Page status | active |
| Metric | Pass@1 accuracy (official); Cons@k majority-vote accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 350 |
| Dataset licence | MIT |
| Publisher | Renmin University of China (Gaoling School of Artificial Intelligence and School of Information), with DataCanvas Alaya NeW and BAAI |
OlymMATH gives a model a high-school Olympiad mathematics problem in English or Chinese. The 2026 paper (arXiv v3) treats 350 unique problems as one suite: 200 computational items with a single numeric or interval answer, plus 150 non-overlapping Lean 4 formalizations for process-level proof checking. The numeric half sits in four expert-labelled fields (algebra, geometry, number theory, combinatorics) and is split into an AIME-level easy 100 and a harder 100. Diagrams were rewritten as text. Numeric answers are restricted so sympy or Math-Verify can grade them without an LLM judge; Lean items must compile.
Free-response: read a text Olympiad problem. Numeric items require a boxed number or interval; Lean items require a compiling Lean 4 proof. Official numeric evaluation uses Pass@1 (mean accuracy over samples) and Cons@k (majority vote). OpenCompass instead grades numeric items with an LLM judge.
No model card in ModelSpec reports this benchmark yet.