SciEval

About 18,000 objective and subjective questions in chemistry, physics and biology, scored on basic knowledge, application, calculation and research ability.

Also known as: SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryscientific knowledge and reasoning across chemistry, physics and biology
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
PublisherShanghai Jiao Tong University (X-LANCE Lab / OpenDFM)

What it measures

SciEval tests scientific understanding across chemistry, physics and biology. Questions are organised by Bloom's-taxonomy-inspired ability levels: basic knowledge (recall of facts and definitions), knowledge application (using facts in a new context), scientific calculation (numeric problem solving) and research ability (experimental design and interpretation). Most items are multiple choice or true/false; a smaller share are open-ended judgement or fill-in questions. A "dynamic" chemistry and physics subset regenerates numeric values from templates at evaluation time so the specific numbers cannot have been memorised from a fixed test file.

Task format

Multiple choice (four options) is the dominant format via the OpenCompass integration, which reads an `input` stem with `A`-`D` options and expects a single letter answer. The original repository's `scieval-test.json` / `scieval-valid.json` also carry true/false and open fill-in-the-blank items scored by the authors' own `eval.py`, and a separate `eval_dynamic.py` regenerates and scores the templated chemistry/physics subset.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub