CompassBench v1.3

OpenCompass's next dated CompassBench release: a narrower, restructured composite of knowledge, math, code and agent scores, not a superset of CompassBench v1.1.

Also known as: compassbench_v1_3

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategoryOpenCompass's own composite evaluation aggregating four category groups -- knowledge, mathematics, code, and agent/tool-use -- restructured from, and not a superset of, CompassBench v1.1
Page statusunknown
Metricfour separate category-level naive averages (knowledge, math, code, agent) -- no single blended top-level score is defined in the shipped configuration, unlike compassbench_20_v1_1
Directionhigher_is_better
Unit%
PublisherOpenCompass (Shanghai AI Laboratory)

What it measures

CompassBench v1.3 is OpenCompass's next dated release in the CompassBench line after `compassbench_20_v1_1`, established directly from its dataset and summarizer configuration files in the opencompass repository, but the evidence shows it is a restructured, genuinely different composite rather than an extended version of v1.1: it aggregates only four category groups -- knowledge (English and Chinese Wikipedia-derived single-choice questions split by domain: humanities, social science, and natural science in both an engineering and a science variant; a fifth domain, life common sense, is defined in the dataset configuration but is not included in the shipped aggregation), mathematics (college-level single choice in English and Chinese, plus an English arithmetic cloze set -- narrower than v1.1's MathBench, which spans arithmetic through college in both languages), code (a HumanEval-style code-completion pair, an LCBench-style code-interview task at three difficulties in both languages, and TACO-based code-competition at five difficulties), and agent (T-Eval / plugin_eval tool-planning only; a CIBench code-interpreter group is defined in the same configuration file but is commented out of the active aggregation). v1.1's language and reasoning (ReasonBench) categories have no counterpart in v1.3 at all. OpenCompass's own release notes separately reference a "CompassBench Checklist Evaluation" and a "Compassbench v1_3 subjective evaluation" landing in the same development window; this page documents only the objective composite under `configs/datasets/compassbench_v1_3`, and does not claim those other, differently-scoped CompassBench-branded evaluations are the same thing.

Task format

Predominantly multiple choice, scored with OpenCompass's circular evaluation (every answer-order permutation must be answered correctly to credit the item, `perf_4`), the same technique v1.1 uses. The knowledge and math categories use this circular multiple-choice format with a step-by-step reasoning instruction in the prompt; the arithmetic-cloze math task and the code tasks are graded by final-answer extraction or functional test execution (pass@1) instead. Unlike v1.1's published summarizer, v1.3's shipped configuration defines no single top-level blended score across all four categories -- the corresponding line in its summarizer file is commented out -- so it is reported as four separate category averages rather than one number.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub