CompassBench v1.1 (public)

The publicly redistributed, reduced-item release of CompassBench v1.1: the same six-category structure, with some code tasks restricted to their first five test cases.

Also known as: compassbench_20_v1_1_public

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategorypublicly redistributed, item-reduced release of CompassBench v1.1's six-category composite
Page statusunknown
Metricsame top-level 'average' construction as compassbench_20_v1_1, computed over the public item set
Directionhigher_is_better
Unit%
PublisherOpenCompass (Shanghai AI Laboratory)

What it measures

Structurally identical to `compassbench_20_v1_1`: the same six category groups (language, knowledge, reasoning, math, code, agent) built from the same dataset-configuration code, confirmed by diffing the two ids' configuration files directly. Every difference found is mechanical rather than substantive -- dataset abbreviations gain a `_public` suffix and read from a separate `data/compassbench_v1.1.public/` directory instead of `data/compassbench_v1.1/` -- except for three code datasets (`mbpp_cn`, `sanitized_mbpp`, `TACO`), where the public configuration adds `test_range='[0:5]'`, restricting evaluation to each item's first five test cases instead of the full set. This id exists so a score can be reproduced outside OpenCompass's own leaderboard infrastructure without necessarily exposing the full item set those internal runs use.

Task format

Identical task formats to compassbench_20_v1_1 (circular multiple choice, BLEU/ROUGE generation, pass@1 code execution); see that page.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub