EESE (Ever-Evolving Science Exam)

A resampled science exam of about 500 items drawn from a 100K+ pool; OpenCompass grades model answers with an LLM judge on a 0–10 scale.

Also known as: Ever-Evolving Science Exam, EESE-V1, EESE-V2, EESE-V3

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorydynamic science exam (closed- and open-ended) across five disciplines
Page statusactive
Metricoverall_score (OpenCompass LLM-judge mean of 0-10 item scores, scaled to 0-100)
Directionhigher_is_better
Unit%
Dataset size486
Dataset licenceMIT
PublisherShanghai Artificial Intelligence Laboratory (AIBENCH / aiben.ch)

What it measures

EESE tests scientific question answering across five disciplines: agricultural sciences, natural sciences, engineering and technology, medical sciences, and humanities and social sciences. Items mix closed-ended forms (single-choice, multiple-choice, fill-in, true/false) with open-ended problems. The evaluated set is a small, periodically resampled slice of EESE-Pool, a larger expert-built repository of more than 100,000 question–answer pairs spanning 500-plus subfields. English text. The skill is scientific knowledge and problem solving, not a single exam subject.

Task format

Zero-shot generation. OpenCompass prompts with the question and question_type: closed items should emit option letters only; open items should show working. A separate LLM judge scores the prediction against the gold final_answer.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub