A 4,428-problem olympiad-level mathematics benchmark built after GSM8K and MATH became easy for frontier models, graded by an LLM judge rather than exact string match.
unassessed
| Category | math |
|---|---|
| Subcategory | olympiad mathematics |
| Page status | active |
| Metric | accuracy (LLM-judged answer equivalence) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 4428 |
| Dataset licence | Apache-2.0 |
| Publisher | Peking University; Alibaba |
Omni-MATH gives a model an olympiad-level mathematics competition problem -- drawn from international, national and regional competitions such as the IMO, Putnam, USAMO and HMMT -- and asks for a full solution and final answer, with no answer choices. Problems are categorised into 33-plus sub-domains (algebra, number theory, geometry, combinatorics, calculus and more) and assigned one of ten difficulty levels, calibrated mainly against the Art of Problem Solving community's own difficulty ratings. It targets reasoning clearly beyond grade-school or standard competition mathematics: the authors built it specifically because GSM8K and the original MATH dataset were, by their account, already being solved with high accuracy by 2024-era models.
Free-response: read an olympiad-level mathematics problem, produce a full solution and final answer; graded by comparing the extracted final answer to a reference answer.
No model card in ModelSpec reports this benchmark yet.