{
 "body": "\n## What it measures\n\nIPhO 2025 Theory evaluates whether a model can solve the same three multi-part theoretical physics\nproblems set for the 56th International Physics Olympiad, held at \u00c9cole Polytechnique in Palaiseau,\nnear Paris, France, from 18 to 24 July 2025. Each problem requires deriving a result from physical\nprinciples \u2014 setting up the right equations and carrying the algebra through, not recalling a fact\nor a formula name \u2014 under exam conditions equivalent to those the roughly 400 human student\ncontestants from 87 teams faced. The task is single-turn and text-based.\n\nThis page classifies IPhO 2025 Theory as a reasoning benchmark rather than a domain-knowledge one,\nbecause solving an olympiad physics problem is closer to the multi-step derivation demanded by\ncompetition-mathematics benchmarks than to closed-book fact recall; the physics content is graduate-\nadjacent but secondary-school-accessible, and the difficulty is almost entirely in the reasoning\nchain rather than in specialised vocabulary.\n\n## How it is scored\n\nThe official theoretical examination is marked out of 30 points across its three problems, over a\n5-hour sitting. For the human contestants, scripts are marked in the writing style required by the\ngeneral instructions \u2014 concise, using equations and symbols rather than prose, since markers are\nnot assumed to be multilingual \u2014 against an official mark scheme, then moderated between each\nteam's leader and the local markers. When a genuine ambiguity was found in one 2025 problem, the\nInternational Board's minutes record that the resolution was to mark answers based on internal\nconsistency with the rest of a student's own work, rather than reworking the scheme. Reported model\npercentages appear to be a raw score normalised against the 30-point maximum, but the exact\nnormalisation and grading process behind any single published number was not established here.\n\n## Dataset and licence\n\nThe theoretical examination comprises three problems, worth 30 points combined, plus one unused\nbackup problem (on strongly correlated fermionic matter) that was shown to team leaders after the\ncompetition but not administered. Official problem sheets, answer sheets and a general data sheet of\nphysical constants were published in English, the reference language, on the organisers' website;\nhuman contestants sat the exam in their own language via officially prepared translations. No\nexplicit licence statement for the exam materials was found in the sources read.\n\n## Who publishes it\n\nThe International Physics Olympiad is organised annually by a rotating host country under the\noversight of an International Board; the 2025 edition was hosted by France, with the theoretical\nexamination held on 21 July 2025 and the results, moderation process and medal thresholds finalised\nby the International Board on 23 July 2025 under President Rajdeep Singh Rawat. IPhO itself dates\nto 1967, per the organisers' own description of the competition. No specific problem-setting\ncommittee or individual author list for the 2025 papers was found in the sources read.\n\n## Lineage\n\nThis appears to be the only IPhO-related page in this repository; no page exists yet for other IPhO\nyears, for the paired experimental examination (two problems, also part of the 2025 competition but\nscored separately and not covered by this page), or for a general \"IPhO\" family page. Each year's\nOlympiad produces a fresh set of problems under the same format, so a natural successor\n(e.g., an IPhO 2026 theory page, for the edition scheduled to be hosted by Colombia) would exist once\nthat year's exam and any independent scoring of models against it are established, but no such page\nexists here yet.\n\n## Saturation and contamination\n\nThis repository found no independent, multi-model tracker scoring current models against the\nofficial IPhO 2025 mark scheme the way, for instance, Epoch AI tracks GPQA Diamond, so saturation\ncannot be assessed here; any single model's reported score should be read as a self-reported,\nin-house evaluation until shown otherwise. Contamination risk sits at medium: the official questions\nand answer sheets have been public on the organisers' site since shortly after the exam, so a model\ntrained on data gathered after July 2025 \u2014 including ordinary web crawls, given the exam's press\ncoverage \u2014 has a real chance of having encountered the problems and their solutions, which weakens\nany claim that a later-trained model is solving them from first principles rather than partial\nrecall.\n\n## How to run it\n\nNo standard automated evaluation harness (lm-evaluation-harness, inspect_evals, HELM, OpenCompass,\nBIG-bench) was confirmed to include this exam. For the actual human competition, grading follows a\nspecific, documented process: the host country's markers grade every script against the official\nmark scheme; each team's leader may request moderation directly with those markers; a three-person\nInternational Board panel (Jevgenij Chmeliov, Andrzej Kotlicki and Stefan Petersen for 2025) hears\ndisputed moderations; and the full International Board votes to approve the final marks before medal\nthresholds are set. No comparably independent or moderated process was found for AI-model\nevaluations of this exam: publishers reporting a model's score appear, from what could be\nestablished here, to grade the model's own written output against the public mark scheme in-house,\nwithout the cross-checking a human contestant's script receives. That gap is worth stating plainly\nwhenever a score for this benchmark is quoted, since olympiad grading is normally exactly the kind\nof process that catches partial credit, notation choices and borderline answers that an automated or\nsingle-grader check might not.\n\n## Reading the numbers\n\nA high score on IPhO 2025 Theory suggests a model can carry out the kind of multi-step, from-first-\nprinciples physics derivation that separates strong human contestants from average ones, on\nproblems written for exactly that purpose. It does not establish that the score was produced under\ngrading as rigorous as the human competition's own moderated, multi-party process, since this\nrepository found no evidence that AI evaluations of this exam are graded and cross-checked the way\nstudent scripts are. Because the questions and official solutions are now public, treat scores from\nmodels with training data extending past July 2025 with real caution, and prefer, where available, a\nreport that states who graded the model's output and how, over a bare percentage.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "ipho_2025_theory",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Muse Spark",
   "model_id": "meta/muse-spark",
   "provider": "meta",
   "provider_display": "Meta",
   "score": 82.6,
   "source": "meta-blog, officechai, artificial-analysis"
  }
 ],
 "page": {
  "aliases": [
   "International Physics Olympiad 2025 Theoretical Examination",
   "IPhO 2025 Theoretical Exam"
  ],
  "category": "reasoning",
  "contamination": {
   "note": "The official problem sheets and answer sheets were published on the organisers' website shortly after the exam, so the exact questions and official solutions have been public since July 2025. A model trained on data collected after that date, including general web crawls, has a real chance of having seen them, similar to the pattern for other olympiad-style LLM benchmarks; scores are most informative for models with a training cutoff before the exam, or evaluated soon after it, before circulation is wide.\n",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": 3,
   "size_note": "Three theoretical problems, worth 30 points combined, sat over 5 hours on 21 July 2025. A fourth \"backup\" theory problem, on strongly correlated fermionic matter, was prepared but not used and was shown to team leaders after the exam.\n",
   "splits": "single exam instance; no train/test split",
   "url": "https://www.ipho2025.fr/sujets-officiels-ipho-france-2025"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "sonnet-5 agent, batch 1b, slice P",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "No standard automated harness was confirmed to include this exam. For the human competition, scripts are marked by the host country's markers against the official mark scheme, then each team's leader can request moderation with the markers; a three-person International Board panel (for 2025: Jevgenij Chmeliov, Andrzej Kotlicki and Stefan Petersen) resolves disputed moderations, and the full International Board votes to approve the final marks before medal thresholds are set \u2014 a process documented in the official IPhO 2025 minutes. This repository found no evidence of an equivalent independent, moderated process for grading AI models: where publishers report a model's score against this exam, they appear, as far as could be established here, to grade the model's output in-house against the public mark scheme, which is a materially different and less independently checked process than how the human contestants' scripts were scored.\n"
  },
  "id": "ipho_2025_theory",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "IPhO 2025 Theory evaluates whether a model can solve the same three multi-part theoretical physics problems that competed for medals at the 56th International Physics Olympiad, held in France in July 2025. Each problem requires deriving results algebraically from first principles and physical reasoning \u2014 not recalling a fact \u2014 under the same official mark scheme used for the roughly 400 human student contestants. It is single-turn, text-based (this repository did not confirm whether the official problem sheets also include diagrams), and testing is closer to multi-step quantitative reasoning than to closed-book knowledge recall, which is why this page treats it as a reasoning benchmark rather than a domain-knowledge one.\n",
  "metric": {
   "baseline_note": "The official theoretical examination is marked out of 30 points across the three problems (per the organisers' general instructions document); reported percentages for models appear to be normalisations of a raw score against that 30-point maximum, though the exact convention used by any given reporter is not established here. The IPhO 2025 medal thresholds published by the International Board (gold 31.1, silver 22.9, bronze 15.3, honourable mention 11.3) are combined theory-plus-experimental totals, not theory-only, so they should not be read as a theory-only human baseline; a theory-only human baseline was not established from the sources read.\n",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "percentage of maximum theory score",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "IPhO 2025 Theory",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "",
   "title": "",
   "url": "",
   "year": null
  },
  "publisher": {
   "authors": [],
   "org": "International Physics Olympiad, 2025 edition, hosted by France (Palaiseau, \u00c9cole Polytechnique, near Paris)",
   "url": "https://www.ipho2025.fr/"
  },
  "released": "2025-07",
  "repo_url": "",
  "saturation": {
   "as_of": "",
   "note": "This repository found no independent tracker that scores multiple frontier models against the official IPhO 2025 marking scheme under one stated protocol (the way, for example, Epoch AI tracks GPQA Diamond), so saturation cannot be assessed here. Scores that circulate for individual models trace back to the model publishers' own announcements rather than to a neutral third party.\n",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "IPhO 2025 France \u2014 official site",
    "url": "https://www.ipho2025.fr/"
   },
   {
    "accessed": "2026-09-08",
    "title": "Sujets officiels (official exam papers), IPhO 2025 France",
    "url": "https://www.ipho2025.fr/sujets-officiels-ipho-france-2025"
   },
   {
    "accessed": "2026-09-08",
    "title": "Theory G0: General instructions, Theoretical Examination (30 points), IPhO 2025",
    "url": "https://cdn.prod.website-files.com/664df830da8ff5d22656764b/687e51b3619a60155a5a2549_exam-theory-G0-english_2025-07-20_1916-UTC.pdf"
   },
   {
    "accessed": "2026-09-08",
    "title": "International Physics Olympiad \u2014 official organisation site",
    "url": "https://www.ipho-new.org/"
   },
   {
    "accessed": "2026-09-08",
    "title": "Minutes of IPhO 2025",
    "url": "https://www.ipho-new.org/physics/wp-content/uploads/2025/08/Minutes-IPhO-2025-France.pdf"
   }
  ],
  "status": "active",
  "subcategory": "olympiad physics, theoretical examination",
  "summary": "The three-problem, 30-point theoretical examination of the 2025 International Physics Olympiad, used to test models against a fresh, human-graded physics exam.",
  "tags": [
   "physics",
   "olympiad",
   "exam",
   "single-instance",
   "human-graded"
  ],
  "task_format": "Three multi-part theoretical physics problems, each requiring worked derivations and numerical or symbolic answers; a model produces written solutions to be graded against the official IPhO 2025 marking scheme, the same instrument used for human contestants' scripts.\n"
 }
}