A five-level scientific-knowledge benchmark, from memorization to real-world application, across biology, chemistry, physics and materials science, with two incompatible dataset versions in current use.
unassessed
| Category | domain |
|---|---|
| Subcategory | multi-level scientific-knowledge evaluation across biology, chemistry, physics and materials science |
| Page status | active |
| Metric | task-dependent: exact match, model-graded fact/TF/MCQ/rating judgments, BLEU/ROUGE average, SMILES molecule similarity, Smith-Waterman protein alignment, or relation-extraction F1 |
| Direction | higher_is_better |
| Dataset size | 28392 |
| Dataset licence | MIT (Hugging Face dataset card); no separate licence file was found in the reference GitHub repository |
| Publisher | Zhejiang University (College of Computer Science and Technology; ZJU-Hangzhou Global Scientific and Technological Innovation Center) |
SciKnowEval tests scientific knowledge and skill across five progressive levels, named after stages from the Confucian "Doctrine of the Mean": Studying Extensively (L1, memorizing facts and literature), Enquiring Earnestly (L2, comprehension -- summarizing, extracting relations, verifying hypotheses), Thinking Profoundly (L3, calculation and multi-step reasoning), Discerning Clearly (L4, safety judgment and harmful-request refusal) and Practicing Assiduously (L5, open-ended application such as designing an experimental protocol or a candidate molecule). Each level is instantiated separately across four scientific domains -- biology, chemistry, physics and materials science -- for a total of dozens of distinct tasks; for example, L1 chemistry includes molecular name conversion and literature QA, while L5 chemistry includes molecular generation and reagent-design tasks. The five-level structure is the benchmark's organising idea: rather than one flat accuracy number, it profiles where in this memory-to-application progression a model's scientific competence breaks down.
A heterogeneous mix of true/false, multiple-choice, relation-extraction and free-form generation tasks (including molecule and protein-sequence generation), evaluated zero-shot, English only, filterable by domain, level or individual task.
No model card in ModelSpec reports this benchmark yet.