R-Bench (Reasoning Bench)

A graduate-level, English/Chinese multi-disciplinary reasoning benchmark - 1,094 text questions across 108 subjects and 665 multimodal questions across 83 subjects, calibrated for Olympiad-level difficulty.

Also known as: RBench, Reasoning Bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorygraduate-level multi-disciplinary reasoning (text and multimodal)
Page statusactive
Metricaccuracy (top-1, single-letter answer)
Directionhigher_is_better
Unit%
Dataset size3518
Dataset licenceApache 2.0
PublisherTsinghua University, with Stanford University, Carnegie Mellon University, University of Pennsylvania, Tencent Hunyuan X and Fitten

What it measures

R-Bench (the authors expand "R" as Reasoning) tests complex, graduate-level reasoning across many academic and professional disciplines rather than a single subject. It has two tracks: RBench-T, text-only multiple-choice questions spanning 108 subjects across 19 departments (mathematics, physics, biology, computer science, chemistry, and more, plus applied areas such as law, finance and medicine), and RBench-M, a multimodal track of 665 image-plus-text questions across 83 subjects. Both tracks are released in parallel English and Chinese versions with the same questions, so cross-lingual consistency can be checked directly. Questions are calibrated for difficulty and subject balance and are intended to be Olympiad-level rather than textbook recall.

Task format

Multiple-choice questions with up to six options (A-F); RBench-M questions additionally include one or more images. Models are prompted zero-shot to answer and asked to end their response with "ANSWER: $LETTER"; OpenCompass ships both a direct letter-extraction scorer and an LLM-judge scorer that compares a free-form response against the gold answer for cases where letter extraction is unreliable.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub