SciKnowEval

A five-level scientific-knowledge benchmark, from memorization to real-world application, across biology, chemistry, physics and materials science, with two incompatible dataset versions in current use.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorymulti-level scientific-knowledge evaluation across biology, chemistry, physics and materials science
Page statusactive
Metrictask-dependent: exact match, model-graded fact/TF/MCQ/rating judgments, BLEU/ROUGE average, SMILES molecule similarity, Smith-Waterman protein alignment, or relation-extraction F1
Directionhigher_is_better
Dataset size28392
Dataset licenceMIT (Hugging Face dataset card); no separate licence file was found in the reference GitHub repository
PublisherZhejiang University (College of Computer Science and Technology; ZJU-Hangzhou Global Scientific and Technological Innovation Center)

What it measures

SciKnowEval tests scientific knowledge and skill across five progressive levels, named after stages from the Confucian "Doctrine of the Mean": Studying Extensively (L1, memorizing facts and literature), Enquiring Earnestly (L2, comprehension -- summarizing, extracting relations, verifying hypotheses), Thinking Profoundly (L3, calculation and multi-step reasoning), Discerning Clearly (L4, safety judgment and harmful-request refusal) and Practicing Assiduously (L5, open-ended application such as designing an experimental protocol or a candidate molecule). Each level is instantiated separately across four scientific domains -- biology, chemistry, physics and materials science -- for a total of dozens of distinct tasks; for example, L1 chemistry includes molecular name conversion and literature QA, while L5 chemistry includes molecular generation and reagent-design tasks. The five-level structure is the benchmark's organising idea: rather than one flat accuracy number, it profiles where in this memory-to-application progression a model's scientific competence breaks down.

Task format

A heterogeneous mix of true/false, multiple-choice, relation-extraction and free-form generation tasks (including molecule and protein-sequence generation), evaluated zero-shot, English only, filterable by domain, level or individual task.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub