An automated chemistry benchmark of nearly 2,800 question-answer pairs, built specifically to compare frontier LLMs against surveyed human chemists rather than just each other.
unassessed
| Category | domain |
|---|---|
| Subcategory | chemistry knowledge and reasoning |
| Page status | active |
| Metric | accuracy (fraction of questions answered correctly) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2788 |
| Dataset licence | MIT |
| Publisher | Friedrich Schiller University Jena (Laboratory of Organic and Macromolecular Chemistry); Helmholtz Institute for Polymers in Energy Applications Jena (HIPOLE Jena); a multi-institution author collaboration |
ChemBench evaluates chemical knowledge and reasoning: questions range from general, inorganic, analytical and technical chemistry to toxicity and safety, and are separately tagged by which skills they require -- knowledge, reasoning, calculation, or chemical intuition -- and by difficulty (basic or advanced). Molecules are encoded as SMILES strings wrapped in explicit START_SMILES/END_SMILES tags so the format is unambiguous to a text-only model. The benchmark's distinguishing feature is that it was built specifically to be comparable to human performance: a curated subset was also given to surveyed chemists so model scores can be read against a same-questions human baseline rather than an assumed one.
A mix of multiple-choice (2,544 questions) and open-ended free-response questions (244 questions), evaluated on text completions only, including from tool-augmented systems. A curated 236-question subset, ChemBench-Mini, restricted to "advanced"-difficulty items and balanced across topic/skill combinations, was answered by human volunteers through a custom web application for the paper's human-baseline comparison; some of those questions allowed the volunteers to use external tools such as web search.
No model card in ModelSpec reports this benchmark yet.