Chem Exam

An OpenCompass-only chemistry benchmark of gaokao-style exam and competition problems, LLM-judge scored for partial credit; no paper, author or publicly locatable dataset was found.

Also known as: Chem_exam

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorychemistry exam and competition problem solving, LLM-judge partial credit
Page statusunknown
MetricLLM-judge partial-credit score (proportion of correctly answered sub-questions)
Directionhigher_is_better
Unit%
Dataset licenceNot established for the question data itself; the OpenCompass repository hosting the evaluation configuration is Apache License 2.0, but that licenses the harness code, not necessarily the underlying chemistry questions.

What it measures

Chem Exam tests chemistry problem-solving across two distinct question sources: a "gaokao" subset modelled on China's national college-entrance exam, and a "competition" subset drawn from chemistry-olympiad-style contest problems. Both subsets present a chemistry question, sometimes containing several numbered sub-questions and sub-sub-questions of increasing specificity, and ask the model to reason step by step to a final boxed answer. Some items carry a `has_img` field, implying the source question included an image such as a molecular diagram or reaction scheme; the generation configuration this page reviewed forwards only the question's text to the model, so it is unclear whether image content reaches the model through a separate pipeline or is effectively dropped for those items. It is a single-turn, free-response, text-primary task. No author, paper, or publisher statement describing how either subset was built was found for this page.

Task format

Free-text chemistry question answered with step-by-step reasoning, with the final answer boxed in LaTeX `\boxed{}` notation. Each subset also has a "rawprompt" variant that swaps in a raw prompt template instead of the standard instruction-formatted one, apparently intended for base or completion-style models rather than instruction-tuned ones.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub