PHYBench

500 original text-only physics problems scored by expression-tree edit distance; Gemini 2.5 Pro reached 36.9% accuracy versus a 61.9% human baseline.

Also known as: PHYBench, PhyBench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryoriginal text-only physics problems with symbolic answers
Page statusactive
Metricaccuracy and EED Score (Expression Edit Distance)
Directionhigher_is_better
Unit% accuracy; EED points 0-100
Dataset size500
Dataset licenceMIT
PublisherSchool of Physics, Peking University (with Institute for Artificial Intelligence, Peking University, and Beijing Computational Science Research Center)

What it measures

PHYBench asks a model to derive a single symbolic expression for a stated physical quantity from a text-only scenario. Problems span mechanics, electromagnetism, thermodynamics, optics, modern physics, and advanced physics, from high school through Physics Olympiad difficulty. Authors at Peking University wrote original items and filtered them so each has one unambiguous symbolic answer and needs no figure. The benchmark is aimed at physical perception and multi-step reasoning, not formula lookup.

Task format

Free-response: read a physics problem, reason, and box one LaTeX expression. Equivalent algebraic forms count as correct. Equations and decimal approximations are rejected. Official metrics are binary accuracy and EED Score (0-100) from tree edit distance on simplified SymPy expression trees.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub