Earth-Silver

EarthSE's harder 1,000-item Earth-science QA split; OpenCompass scores the 250 multiple-choice items, not Iron, Gold, or the other three formats.

Also known as: Earth_Silver, earth_silver_mcq, EarthSE Silver

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryhard Earth-science QA from high-impact papers (EarthSE middle tier)
Page statusactive
Metricaccuracy (OpenCompass AccEvaluator on MC; paper also reports TF/FIB accuracy and FR win-rate)
Directionhigher_is_better
Unit%
Dataset size1000
Dataset licenceMIT
PublisherShanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University

What it measures

Earth-Silver is the professional-difficulty QA tier of EarthSE (Xu et al.). Items are English questions built from high-impact Earth-science papers, covering five spheres (atmosphere, biosphere, hydrosphere, lithosphere, cryosphere), many sub-disciplines, and 11 task types such as analysis, calculation, experiment design and code generation. Hugging Face ships four formats of 250 items each: multiple_choice, true_false, fill_in_the_blank and free_form. This id, as OpenCompass implements it, loads only split multiple_choice. The skill is hard Earth-science answering, not [climaqa](climaqa.md) climate QA and not Earth-Gold dialogue.

Task format

OpenCompass: zero-shot generation. Prompt "Q: {question}" then "output only the corresponding letter". Earth_Silver_gen.py scores AccEvaluator and mentions options A-D. LLM-judge configs mention A-E and GenericLLMEvaluator. Few-shot is NotImplemented in Earth_Silver_MCQDataset.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub