A graduate-level, English/Chinese multi-disciplinary reasoning benchmark - 1,094 text questions across 108 subjects and 665 multimodal questions across 83 subjects, calibrated for Olympiad-level difficulty.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | graduate-level multi-disciplinary reasoning (text and multimodal) |
| Page status | active |
| Metric | accuracy (top-1, single-letter answer) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 3518 |
| Dataset licence | Apache 2.0 |
| Publisher | Tsinghua University, with Stanford University, Carnegie Mellon University, University of Pennsylvania, Tencent Hunyuan X and Fitten |
R-Bench (the authors expand "R" as Reasoning) tests complex, graduate-level reasoning across many academic and professional disciplines rather than a single subject. It has two tracks: RBench-T, text-only multiple-choice questions spanning 108 subjects across 19 departments (mathematics, physics, biology, computer science, chemistry, and more, plus applied areas such as law, finance and medicine), and RBench-M, a multimodal track of 665 image-plus-text questions across 83 subjects. Both tracks are released in parallel English and Chinese versions with the same questions, so cross-lingual consistency can be checked directly. Questions are calibrated for difficulty and subject balance and are intended to be Olympiad-level rather than textbook recall.
Multiple-choice questions with up to six options (A-F); RBench-M questions additionally include one or more images. Models are prompted zero-shot to answer and asked to end their response with "ANSWER: $LETTER"; OpenCompass ships both a direct letter-extraction scorer and an LLM-judge scorer that compares a free-form response against the gold answer for cases where letter extraction is unreliable.
No model card in ModelSpec reports this benchmark yet.