CritPt (Complex Research using Integrated Thinking — Physics Test)

Seventy public research-physics challenges with held-out answers, graded by a server; full suite is 71 challenges plus 190 checkpoints.

Also known as: CritPt, Critical Point, CritPt_main

active

This benchmark is in the default catalogue: its identity, protocol, current model coverage and dated results were verified by a reviewer who opened the sources.

Recorded reasons:

Categoryreasoning
Subcategoryunpublished research-level physics challenges with machine-verifiable answers
Page statusactive
Metricpass@1 challenge accuracy (mean over 5 runs × 70 test challenges)
Directionhigher_is_better
Unit%
Dataset size70
Dataset licenceApache-2.0
PublisherArgonne National Laboratory, University of Illinois Urbana-Champaign, Virginia Tech, and collaborators

What it measures

CritPt tests whether a model can finish entry-level physics research problems that working scientists wrote from their own unpublished work. Each challenge is meant to look like a junior Ph.D. project: well posed, solvable from public knowledge, and not a contest recitation. Answers are guess-resistant and checked by an automated grader for physics-specific formats. The public Hugging Face dump is the 70-item test set of full challenges. One extra example challenge is shown on the project site. 190 checkpoint subproblems exist in the paper but are not this Hub split.

Task format

Multi-turn markdown physics problem; the model writes a derivation and a "Final Answer:" line, then fills a Python answer() template. OpenCompass default is zero-shot chat with max_out_len 32,768, no code tool and no web search. Official scoring is a full 70-item batch on the Artificial Analysis grading API, not local string match.

Verified results

Each row was checked against its source by a reviewer.

ModelScoreEvidence dateSource kindLink
GPT-6 Astra (max)32.0%2026-09-04 publishedindependent_evaluatorsource
GLM-5.3 (max)19.0%2026-09-04 publishedindependent_evaluatorsource

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub