Seventy public research-physics challenges with held-out answers, graded by a server; full suite is 71 challenges plus 190 checkpoints.
active
Recorded reasons:
| Category | reasoning |
|---|---|
| Subcategory | unpublished research-level physics challenges with machine-verifiable answers |
| Page status | active |
| Metric | pass@1 challenge accuracy (mean over 5 runs × 70 test challenges) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 70 |
| Dataset licence | Apache-2.0 |
| Publisher | Argonne National Laboratory, University of Illinois Urbana-Champaign, Virginia Tech, and collaborators |
CritPt tests whether a model can finish entry-level physics research problems that working scientists wrote from their own unpublished work. Each challenge is meant to look like a junior Ph.D. project: well posed, solvable from public knowledge, and not a contest recitation. Answers are guess-resistant and checked by an automated grader for physics-specific formats. The public Hugging Face dump is the 70-item test set of full challenges. One extra example challenge is shown on the project site. 190 checkpoint subproblems exist in the paper but are not this Hub split.
Multi-turn markdown physics problem; the model writes a derivation and a "Final Answer:" line, then fills a Python answer() template. OpenCompass default is zero-shot chat with max_out_len 32,768, no code tool and no web search. Official scoring is a full 70-item batch on the Artificial Analysis grading API, not local string match.
Each row was checked against its source by a reviewer.
| Model | Score | Evidence date | Source kind | Link |
|---|---|---|---|---|
| GPT-6 Astra (max) | 32.0% | 2026-09-04 published | independent_evaluator | source |
| GLM-5.3 (max) | 19.0% | 2026-09-04 published | independent_evaluator | source |
No model card in ModelSpec reports this benchmark yet.