798 expert-written STEM problems across seven fields, scored by an LLM judge; GPT-5-High is at 42.9% on the public validation set.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | high-difficulty multidisciplinary scientific problem solving |
| Page status | active |
| Metric | accuracy (LLM-judge); also mG-Pass@2 and mG-Pass@4 |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 798 |
| Publisher | Shanghai AI Laboratory (OpenCompass) |
ATLAS gives a model an original scientific problem and asks for a structured final answer, often with several sub-answers, rather than a single letter choice. Items cover mathematics, physics, chemistry, biology, computer science, earth science and materials science. The paper says most items are calculation and derivation, with smaller shares of selection, explanation and composite formats. Problems were written or substantially rewritten by PhD-level domain experts so that the skill is multi-step scientific reasoning, not recall of a public exam item.
Free-form generation. OpenCompass asks the model to solve the problem, then emit a JSON list of final answers. An LLM judge compares that list to a standard answer. Default config draws four samples (n=4) for accuracy and mG-Pass@{2,4}.
No model card in ModelSpec reports this benchmark yet.