AIME 2024

The 30 problems from the 2024 American Invitational Mathematics Examination, an exact-answer competition-math test now well past its useful ceiling for frontier models.

Also known as: AIME24, AIME 2024 I and II

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategorycompetition mathematics
Page statussaturated
Metricaccuracy (pass@1, exact match on the final integer)
Directionhigher_is_better
Unit%
Dataset size30
Dataset licenceApache-2.0
PublisherMathematical Association of America (MAA)

What it measures

AIME 2024 gives a model the problems from the 2024 American Invitational Mathematics Examination, a competition only the top-scoring AMC 10/12 participants are invited to sit, and checks whether the model returns the single correct integer answer. It exercises multi-step algebra, geometry, number theory and combinatorics reasoning well above grade-school math benchmarks, and gives no partial credit for a sound method that lands on the wrong final number. As one of the earliest AIME sittings widely adopted as an LLM benchmark, it is also the sitting most often cited as an example of how quickly a fixed competition-math set can go from differentiating to saturated once reasoning-focused models arrive.

Task format

Free-response competition math problem in, single integer answer from 0 to 999 out; no answer choices are offered.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub