LiveMathBench

A bilingual, periodically re-released competition-math benchmark from recent AMC, CNMO, CCEE and Putnam problems, paired with G-Pass@k, a metric scoring correctness and stability across samples.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategorycontinuously-updated, contamination-resistant competition mathematics
Page statusactive
MetricG-Pass@k and mG-Pass@k (stability-aware pass rate across repeated samples), alongside plain Greedy accuracy
Directionhigher_is_better
Unit%
Dataset size238
Dataset licenceListed as CC BY 4.0 on the dataset card, but the repository is access-gated and requires agreeing to a checkbox reading "I agree to use this dataset for non-commercial use ONLY" before download -- a stated licence and a stated use restriction that are in tension with each other. Both readings are given here rather than picking one.
PublisherShanghai Artificial Intelligence Laboratory

What it measures

LiveMathBench gives a model a recent competition mathematics problem, converted to free-response form (multiple-choice options are stripped so the model must derive the answer itself), drawn from four sources: the China National Mathematical Olympiad (CNMO), China's College Entrance Examination (CCEE, from mock exams), the American Mathematics Competition (AMC), and the William Lowell Putnam Mathematical Competition (WLPMC). Its purpose-built companion metric, G-Pass@k, measures not just whether a model can solve a problem once but whether it solves it consistently across many independent samples, addressing what the authors call a gap between a model's "potential" (can it ever get this right) and its "stability" (does it reliably get this right). A harder subset, LiveMathBench-Hard, is drawn from the same four sources at greater difficulty.

Task format

Free-response: read a competition mathematics problem (in Chinese or English for the v202412 release; English only for v202505) and produce a worked solution ending in a final answer, with no answer choices offered even where the original competition question was multiple choice.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub