8,476 olympiad-level bilingual math and physics problems, many with figures, from contests and the Chinese gaokao; GPT-4V scored 17.97% on the full set at release.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | bilingual olympiad math and physics, text and figures |
| Page status | active |
| Metric | accuracy (open-ended items; automated answer match) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 8476 |
| Dataset licence | MIT on the GitHub LICENSE file; Apache-2.0 on the Hugging Face dataset card (both opened; not reconciled) |
| Publisher | OpenBMB / Tsinghua University |
OlympiadBench asks a model to solve contest mathematics or physics problems at Olympiad and Chinese college-entrance (gaokao) difficulty, in English or Chinese, with or without accompanying figures. Items are open-ended questions or theorem proofs. Each record carries expert annotations: subject, language, answer type, subfield, and a step-by-step solution. The benchmark is meant to be harder than saturated school-math sets and to test scientific reasoning in two languages and two modalities, not only English text.
Zero-shot free response. Open-ended items ask for a numeric, expression, or tuple answer that an automated judge compares to a gold final_answer. Proof items are sampled by humans in the original paper, not auto-scored. OpenCompass ships only five text-only open-ended category files, not the full multimodal set.
No model card in ModelSpec reports this benchmark yet.