FrontierScience

OpenAI expert-level science suite: 100 Olympiad short-answer items and 60 Research rubric tasks in physics, chemistry, and biology.

Also known as: openai/frontierscience

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryexpert-level physics, chemistry, and biology; Olympiad short-answer plus Research rubrics
Page statusactive
MetricOlympiad: binary accuracy via model-graded equivalence; Research: pass rate at ≥7/10 rubric points (paper) or mean normalized rubric score (Inspect)
Directionhigher_is_better
Unit%
Dataset size160
Dataset licenceApache-2.0
PublisherOpenAI

What it measures

FrontierScience tests expert-level scientific reasoning in English, text-only LaTeX, across physics, chemistry, and biology. The Olympiad track is short-answer problems at IPhO / IChO / IBO difficulty, written by medalists and coaches, graded by equivalence to a reference expression, number, formula, or phrase. The Research track is open-ended PhD-level subtasks with a 10-point process rubric. OpenAI wrote several hundred questions and released a 160-item gold set (100 Olympiad + 60 Research); the rest is held out to watch contamination. This family page is the combined suite. The Research track also has its own page at [frontierscience_research](frontierscience_research.md).

Task format

Free-text solution. Inspect Evals task inspect_evals/frontierscience loads openai/frontierscience (pinned 25ed67db7da8f4591484e764008ff585544f5a30), detects format from rubric markers in the answer field, and can filter with -T format=olympic or format=research and -T subjects=physics|chemistry|biology. Default shuffle is True. Official paper judging uses GPT-5 at high reasoning effort, no browsing.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub