BeyondAIME

100 newly written competition math problems at or above the difficulty of AIME's hardest five problems, each manually revised to be unique and to resist guessing, with a single verifiable integer answer.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategorycompetition mathematics harder than AIME, integer-answer
Page statusactive
Metricaccuracy (pass@1, verified answer match)
Directionhigher_is_better
Unit%
Dataset size100
Dataset licenceCC0 1.0 (public domain dedication), per the dataset card
PublisherByteDance Seed

What it measures

BeyondAIME gives a model 100 original competition-mathematics problems built specifically to stay hard once benchmarks like AIME stop separating frontier models. The dataset card states the construction principles directly: every problem targets a difficulty at or above AIME's problems 11-15 (conventionally its hardest third), is manually revised to be a unique formulation not findable in standard pre-training corpora, tests reasoning rather than specialised mathematical knowledge beyond the standard university level, and is deliberately reworked to avoid "pseudo-proof" problems where guessing the final answer is much easier than actually solving the problem. Like AIME, every problem has exactly one positive integer answer, chosen specifically to allow unambiguous, fully automated grading rather than free-form proof evaluation.

Task format

Free-response competition math problem in (Markdown with LaTeX), single positive integer answer out, typically boxed; no answer choices are offered.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub