ATLAS (AGI-Oriented Testbed for Logical Application in Science)

798 expert-written STEM problems across seven fields, scored by an LLM judge; GPT-5-High is at 42.9% on the public validation set.

Also known as: ATLAS, AGI-Oriented Testbed for Logical Application in Science

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryhigh-difficulty multidisciplinary scientific problem solving
Page statusactive
Metricaccuracy (LLM-judge); also mG-Pass@2 and mG-Pass@4
Directionhigher_is_better
Unit%
Dataset size798
PublisherShanghai AI Laboratory (OpenCompass)

What it measures

ATLAS gives a model an original scientific problem and asks for a structured final answer, often with several sub-answers, rather than a single letter choice. Items cover mathematics, physics, chemistry, biology, computer science, earth science and materials science. The paper says most items are calculation and derivation, with smaller shares of selection, explanation and composite formats. Problems were written or substantially rewritten by PhD-level domain experts so that the skill is multi-step scientific reasoning, not recall of a public exam item.

Task format

Free-form generation. OpenCompass asks the model to solve the problem, then emit a JSON list of final answers. An LLM judge compares that list to a standard answer. Default config draws four samples (n=4) for accuracy and mG-Pass@{2,4}.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub