The 30 problems from the 2024 American Invitational Mathematics Examination, an exact-answer competition-math test now well past its useful ceiling for frontier models.
unassessed
| Category | math |
|---|---|
| Subcategory | competition mathematics |
| Page status | saturated |
| Metric | accuracy (pass@1, exact match on the final integer) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 30 |
| Dataset licence | Apache-2.0 |
| Publisher | Mathematical Association of America (MAA) |
AIME 2024 gives a model the problems from the 2024 American Invitational Mathematics Examination, a competition only the top-scoring AMC 10/12 participants are invited to sit, and checks whether the model returns the single correct integer answer. It exercises multi-step algebra, geometry, number theory and combinatorics reasoning well above grade-school math benchmarks, and gives no partial credit for a sound method that lands on the wrong final number. As one of the earliest AIME sittings widely adopted as an LLM benchmark, it is also the sitting most often cited as an example of how quickly a fixed competition-math set can go from differentiating to saturated once reasoning-focused models arrive.
Free-response competition math problem in, single integer answer from 0 to 999 out; no answer choices are offered.
No model card in ModelSpec reports this benchmark yet.