The publicly redistributed, reduced-item release of CompassBench v1.1: the same six-category structure, with some code tasks restricted to their first five test cases.
unassessed
| Category | composite |
|---|---|
| Subcategory | publicly redistributed, item-reduced release of CompassBench v1.1's six-category composite |
| Page status | unknown |
| Metric | same top-level 'average' construction as compassbench_20_v1_1, computed over the public item set |
| Direction | higher_is_better |
| Unit | % |
| Publisher | OpenCompass (Shanghai AI Laboratory) |
Structurally identical to `compassbench_20_v1_1`: the same six category groups (language, knowledge, reasoning, math, code, agent) built from the same dataset-configuration code, confirmed by diffing the two ids' configuration files directly. Every difference found is mechanical rather than substantive -- dataset abbreviations gain a `_public` suffix and read from a separate `data/compassbench_v1.1.public/` directory instead of `data/compassbench_v1.1/` -- except for three code datasets (`mbpp_cn`, `sanitized_mbpp`, `TACO`), where the public configuration adds `test_range='[0:5]'`, restricting evaluation to each item's first five test cases instead of the full set. This id exists so a score can be reproduced outside OpenCompass's own leaderboard infrastructure without necessarily exposing the full item set those internal runs use.
Identical task formats to compassbench_20_v1_1 (circular multiple choice, BLEU/ROUGE generation, pass@1 code execution); see that page.
No model card in ModelSpec reports this benchmark yet.