A bilingual, periodically re-released competition-math benchmark from recent AMC, CNMO, CCEE and Putnam problems, paired with G-Pass@k, a metric scoring correctness and stability across samples.
unassessed
| Category | math |
|---|---|
| Subcategory | continuously-updated, contamination-resistant competition mathematics |
| Page status | active |
| Metric | G-Pass@k and mG-Pass@k (stability-aware pass rate across repeated samples), alongside plain Greedy accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 238 |
| Dataset licence | Listed as CC BY 4.0 on the dataset card, but the repository is access-gated and requires agreeing to a checkbox reading "I agree to use this dataset for non-commercial use ONLY" before download -- a stated licence and a stated use restriction that are in tension with each other. Both readings are given here rather than picking one. |
| Publisher | Shanghai Artificial Intelligence Laboratory |
LiveMathBench gives a model a recent competition mathematics problem, converted to free-response form (multiple-choice options are stripped so the model must derive the answer itself), drawn from four sources: the China National Mathematical Olympiad (CNMO), China's College Entrance Examination (CCEE, from mock exams), the American Mathematics Competition (AMC), and the William Lowell Putnam Mathematical Competition (WLPMC). Its purpose-built companion metric, G-Pass@k, measures not just whether a model can solve a problem once but whether it solves it consistently across many independent samples, addressing what the authors call a gap between a model's "potential" (can it ever get this right) and its "stability" (does it reliably get this right). A harder subset, LiveMathBench-Hard, is drawn from the same four sources at greater difficulty.
Free-response: read a competition mathematics problem (in Chinese or English for the v202412 release; English only for v202505) and produce a worked solution ending in a final answer, with no answer choices offered even where the original competition question was multiple choice.
No model card in ModelSpec reports this benchmark yet.