A resampled science exam of about 500 items drawn from a 100K+ pool; OpenCompass grades model answers with an LLM judge on a 0–10 scale.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | dynamic science exam (closed- and open-ended) across five disciplines |
| Page status | active |
| Metric | overall_score (OpenCompass LLM-judge mean of 0-10 item scores, scaled to 0-100) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 486 |
| Dataset licence | MIT |
| Publisher | Shanghai Artificial Intelligence Laboratory (AIBENCH / aiben.ch) |
EESE tests scientific question answering across five disciplines: agricultural sciences, natural sciences, engineering and technology, medical sciences, and humanities and social sciences. Items mix closed-ended forms (single-choice, multiple-choice, fill-in, true/false) with open-ended problems. The evaluated set is a small, periodically resampled slice of EESE-Pool, a larger expert-built repository of more than 100,000 question–answer pairs spanning 500-plus subfields. English text. The skill is scientific knowledge and problem solving, not a single exam subject.
Zero-shot generation. OpenCompass prompts with the question and question_type: closed items should emit option letters only; open items should show working. A separate LLM judge scores the prediction against the gold final_answer.
No model card in ModelSpec reports this benchmark yet.