ChemBench

An automated chemistry benchmark of nearly 2,800 question-answer pairs, built specifically to compare frontier LLMs against surveyed human chemists rather than just each other.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorychemistry knowledge and reasoning
Page statusactive
Metricaccuracy (fraction of questions answered correctly)
Directionhigher_is_better
Unit%
Dataset size2788
Dataset licenceMIT
PublisherFriedrich Schiller University Jena (Laboratory of Organic and Macromolecular Chemistry); Helmholtz Institute for Polymers in Energy Applications Jena (HIPOLE Jena); a multi-institution author collaboration

What it measures

ChemBench evaluates chemical knowledge and reasoning: questions range from general, inorganic, analytical and technical chemistry to toxicity and safety, and are separately tagged by which skills they require -- knowledge, reasoning, calculation, or chemical intuition -- and by difficulty (basic or advanced). Molecules are encoded as SMILES strings wrapped in explicit START_SMILES/END_SMILES tags so the format is unambiguous to a text-only model. The benchmark's distinguishing feature is that it was built specifically to be comparable to human performance: a curated subset was also given to surveyed chemists so model scores can be read against a same-questions human baseline rather than an assumed one.

Task format

A mix of multiple-choice (2,544 questions) and open-ended free-response questions (244 questions), evaluated on text completions only, including from tool-augmented systems. A curated 236-question subset, ChemBench-Mini, restricted to "advanced"-difficulty items and balanced across topic/skill combinations, was answered by human volunteers through a custom web application for the paper's human-baseline comparison; some of those questions allowed the volunteers to use external tools such as web search.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub