OlympiadBench

8,476 olympiad-level bilingual math and physics problems, many with figures, from contests and the Chinese gaokao; GPT-4V scored 17.97% on the full set at release.

Also known as: OlympiadBench, Olympiad Bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategorybilingual olympiad math and physics, text and figures
Page statusactive
Metricaccuracy (open-ended items; automated answer match)
Directionhigher_is_better
Unit%
Dataset size8476
Dataset licenceMIT on the GitHub LICENSE file; Apache-2.0 on the Hugging Face dataset card (both opened; not reconciled)
PublisherOpenBMB / Tsinghua University

What it measures

OlympiadBench asks a model to solve contest mathematics or physics problems at Olympiad and Chinese college-entrance (gaokao) difficulty, in English or Chinese, with or without accompanying figures. Items are open-ended questions or theorem proofs. Each record carries expert annotations: subject, language, answer type, subfield, and a step-by-step solution. The benchmark is meant to be harder than saturated school-math sets and to test scientific reasoning in two languages and two modalities, not only English text.

Task format

Zero-shot free response. Open-ended items ask for a numeric, expression, or tuple answer that an automated judge compares to a gold final_answer. Proof items are sampled by humans in the original paper, not auto-scored. OpenCompass ships only five text-only open-ended category files, not the full multimodal set.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub