Putnam-AXIOM

522 Putnam contest problems scored by boxed exact match, plus 100 functional variants meant to catch memorisation of public contest write-ups.

Also known as: Putnam AXIOM, Putnam-AXIOM Original, putnam_axiom_original

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategoryuniversity Putnam contest problems with boxed answers and functional variants
Page statusactive
Metricexact_match (boxed answer, SymPy equivalence)
Directionhigher_is_better
Unit%
Dataset size522
Dataset licenceApache-2.0
PublisherStanford University, Department of Computer Science

What it measures

Putnam-AXIOM tests whether a model can solve undergraduate contest mathematics that still separates today's systems. Each item is a William Lowell Putnam problem from 1938–2023 rewritten so the solution ends in one boxed numeric or algebraic answer. The original 522-item set is a static exam. A companion variation protocol rewrites 100 of those problems by changing variables, constants, and surface wording while keeping the same reasoning. The paper reports that even strong models drop when the numbers change, which is the point of the variation split: to show scores that may come from having seen the public contest archive rather than from solving a new instance.

Task format

English LaTeX problem statement. The model writes a solution and a final answer inside \\boxed{}. lm-eval uses a 4-shot Minerva-style prompt, greedy generate_until, and a SymPy/LaTeX equivalence check on the boxed string.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub