An OpenCompass-only chemistry benchmark of gaokao-style exam and competition problems, LLM-judge scored for partial credit; no paper, author or publicly locatable dataset was found.
unassessed
| Category | domain |
|---|---|
| Subcategory | chemistry exam and competition problem solving, LLM-judge partial credit |
| Page status | unknown |
| Metric | LLM-judge partial-credit score (proportion of correctly answered sub-questions) |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | Not established for the question data itself; the OpenCompass repository hosting the evaluation configuration is Apache License 2.0, but that licenses the harness code, not necessarily the underlying chemistry questions. |
Chem Exam tests chemistry problem-solving across two distinct question sources: a "gaokao" subset modelled on China's national college-entrance exam, and a "competition" subset drawn from chemistry-olympiad-style contest problems. Both subsets present a chemistry question, sometimes containing several numbered sub-questions and sub-sub-questions of increasing specificity, and ask the model to reason step by step to a final boxed answer. Some items carry a `has_img` field, implying the source question included an image such as a molecular diagram or reaction scheme; the generation configuration this page reviewed forwards only the question's text to the model, so it is unclear whether image content reaches the model through a separate pipeline or is effectively dropped for those items. It is a single-turn, free-response, text-primary task. No author, paper, or publisher statement describing how either subset was built was found for this page.
Free-text chemistry question answered with step-by-step reasoning, with the final answer boxed in LaTeX `\boxed{}` notation. Each subset also has a "rawprompt" variant that swaps in a raw prompt template instead of the standard instruction-formatted one, apparently intended for base or completion-style models rather than instruction-tuned ones.
No model card in ModelSpec reports this benchmark yet.