KCLE (OpenCompass)

OpenCompass multiple-choice set (abbrs kcle and kcle_fix) graded by an LLM judge; item count, licence and the acronym's expansion are not established.

Also known as: KLCE, kcle_fix

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryOpenCompass local multiple-choice set graded by an LLM judge (expansion of KCLE/KLCE not established)
Page statusunknown
MetricLLM-judge accuracy (GenericLLMEvaluator, A=CORRECT / B=INCORRECT)
Directionhigher_is_better
Unit%
PublisherOpenCompass / Shanghai AI Laboratory (harness packaging); original dataset authors not named in the loader

What it measures

kcle is OpenCompass's KCLEDataset: a local jsonl of input/target pairs that the default configs treat as multiple-choice questions and score with an LLM judge. The mapping table comments the block as "KLCE Datasets" while paths and class names use kcle, so the acronym itself is not established from a paper. Local filenames are kcle_diamond.jsonl and kcle_diamond_fix_251029.jsonl. That "diamond" token is not evidence that this id is [gpqa_diamond](gpqa_diamond.md); GPQA Diamond is a separate 198-item science set with its own page. What the items actually ask is not described beyond the generic multiple-choice prompt, because the jsonl files and Hub paths were not readable here.

Task format

Zero-shot generation (ZeroRetriever + GenInferencer). One config feeds the raw input field. The kcle_fix configs prepend an English instruction to end with ANSWER: $LETTER. A second model (GenericLLMEvaluator) then grades the completion against the gold target as A/CORRECT or B/INCORRECT.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub