OpenAI expert-level science suite: 100 Olympiad short-answer items and 60 Research rubric tasks in physics, chemistry, and biology.
unassessed
| Category | domain |
|---|---|
| Subcategory | expert-level physics, chemistry, and biology; Olympiad short-answer plus Research rubrics |
| Page status | active |
| Metric | Olympiad: binary accuracy via model-graded equivalence; Research: pass rate at ≥7/10 rubric points (paper) or mean normalized rubric score (Inspect) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 160 |
| Dataset licence | Apache-2.0 |
| Publisher | OpenAI |
FrontierScience tests expert-level scientific reasoning in English, text-only LaTeX, across physics, chemistry, and biology. The Olympiad track is short-answer problems at IPhO / IChO / IBO difficulty, written by medalists and coaches, graded by equivalence to a reference expression, number, formula, or phrase. The Research track is open-ended PhD-level subtasks with a 10-point process rubric. OpenAI wrote several hundred questions and released a 160-item gold set (100 Olympiad + 60 Research); the rest is held out to watch contamination. This family page is the combined suite. The Research track also has its own page at [frontierscience_research](frontierscience_research.md).
Free-text solution. Inspect Evals task inspect_evals/frontierscience loads openai/frontierscience (pinned 25ed67db7da8f4591484e764008ff585544f5a30), detects format from rubric markers in the answer field, and can filter with -T format=olympic or format=research and -T subjects=physics|chemistry|biology. Default shuffle is True. Official paper judging uses GPT-5 at high reasoning effort, no browsing.
No model card in ModelSpec reports this benchmark yet.