OlymMATH

A 350-problem bilingual Olympiad math suite: 200 numeric easy/hard items plus 150 Lean 4 proofs, after AIME and MATH stopped separating frontier models.

Also known as: OlymMATH, OlymMATH-EN, OlymMATH-ZH, OlymMATH-HARD, OlymMATH-EASY

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategorybilingual olympiad mathematics with numeric answers and Lean proofs
Page statusactive
MetricPass@1 accuracy (official); Cons@k majority-vote accuracy
Directionhigher_is_better
Unit%
Dataset size350
Dataset licenceMIT
PublisherRenmin University of China (Gaoling School of Artificial Intelligence and School of Information), with DataCanvas Alaya NeW and BAAI

What it measures

OlymMATH gives a model a high-school Olympiad mathematics problem in English or Chinese. The 2026 paper (arXiv v3) treats 350 unique problems as one suite: 200 computational items with a single numeric or interval answer, plus 150 non-overlapping Lean 4 formalizations for process-level proof checking. The numeric half sits in four expert-labelled fields (algebra, geometry, number theory, combinatorics) and is split into an AIME-level easy 100 and a harder 100. Diagrams were rewritten as text. Numeric answers are restricted so sympy or Math-Verify can grade them without an LLM judge; Lean items must compile.

Task format

Free-response: read a text Olympiad problem. Numeric items require a boxed number or interval; Lean items require a compiling Lean 4 proof. Official numeric evaluation uses Pass@1 (mean accuracy over samples) and Cons@k (majority vote). OpenCompass instead grades numeric items with an LLM judge.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub