FrontierScience — Research track

60 open-ended, PhD-level scientific research subtasks in physics, chemistry and biology, graded on a 10-point rubric.

Also known as: FrontierScience Research

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryPhD-level scientific research tasks
Page statusactive
Metricpass rate (share of tasks scoring at least 7 of 10 rubric points)
Directionhigher_is_better
Unit%
Dataset size60
Dataset licenceApache-2.0
PublisherOpenAI

What it measures

The Research track of FrontierScience asks a model to work through a self-contained, multi-step subtask at the level of difficulty a PhD scientist might meet during actual research — for example, proposing and justifying a synthesis route, or analysing an experimental result — rather than answering a closed-form question. It spans physics, chemistry and biology, is text-only and English-language, and is explicitly built to test open-ended scientific reasoning where earlier multiple-choice science benchmarks had become saturated.

Task format

A multi-step, open-ended scientific research subtask answered in free text; graded by a model-based grader (GPT-5) against a per-question rubric worth 10 points, with 7 or more counted as correct.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Muse SparkMeta38.32026-04

Data

This page as JSON · Edit on GitHub