Omni-MATH

A 4,428-problem olympiad-level mathematics benchmark built after GSM8K and MATH became easy for frontier models, graded by an LLM judge rather than exact string match.

Also known as: OmniMATH, Omni-Math

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategoryolympiad mathematics
Page statusactive
Metricaccuracy (LLM-judged answer equivalence)
Directionhigher_is_better
Unit%
Dataset size4428
Dataset licenceApache-2.0
PublisherPeking University; Alibaba

What it measures

Omni-MATH gives a model an olympiad-level mathematics competition problem -- drawn from international, national and regional competitions such as the IMO, Putnam, USAMO and HMMT -- and asks for a full solution and final answer, with no answer choices. Problems are categorised into 33-plus sub-domains (algebra, number theory, geometry, combinatorics, calculus and more) and assigned one of ten difficulty levels, calibrated mainly against the Art of Problem Solving community's own difficulty ratings. It targets reasoning clearly beyond grade-school or standard competition mathematics: the authors built it specifically because GSM8K and the original MATH dataset were, by their account, already being solved with high accuracy by 2024-era models.

Task format

Free-response: read an olympiad-level mathematics problem, produce a full solution and final answer; graded by comparing the extracted final answer to a reference answer.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub