{
 "body": "\nPart of the [Humanity's Last Exam](hle.md) family.\n\n## What it measures\n\nThis id is the same 2,500-question HLE set, scored when the model was allowed tools rather than\nanswering closed-book. It is a different measurement from `hle`, not a harder or easier version of it:\na model can look up or compute an answer instead of recalling or deriving it. The exact tool setting\ndiffers by source. Anthropic's Claude Opus 4.5 System Card (November 2025) is the clearest documented\ncase: it defines \"with search\" as web search, web fetch and code execution, run without extended\nthinking and graded by a separate model, with a documented decontamination pass over the transcripts.\nTreat any other source's \"with tools\" HLE number as unverified until its own configuration is known.\n\n## Reading the numbers\n\nTool access consistently raises HLE scores over the same model's no-tools baseline \u2014 Anthropic's\nNovember 2025 figures show roughly an 8 to 13 point gain across six models \u2014 reflecting that some HLE\nquestions are answerable by finding a source online despite the benchmark's intent to resist lookup.\nA high `hle_tools` score is therefore better read as research and tool-use competence than pure\nknowledge, and should be paired with the same model's plain `hle` score to see how much is retrieval. Contamination is a bigger concern here, since live search can retrieve a\nleaked answer directly; check the decontamination method, if any, before trusting a reported number.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "hle_tools",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Mythos Preview",
   "model_id": "anthropic/claude-mythos-preview",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 64.7,
   "source": "anthropic-system-card"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Opus 4.6",
   "model_id": "anthropic/claude-opus-4-6",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 53.1,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, anthropic-system-card-mythos"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GLM 5.1",
   "model_id": "zhipu/glm-5-1",
   "provider": "zhipu",
   "provider_display": "Z.ai (Zhipu AI)",
   "score": 52.3,
   "source": "zai-org-model-card"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.4",
   "model_id": "openai/gpt-5-4",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 52.1,
   "source": "lmarena.ai, provider-reports, anthropic-system-card-mythos, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Gemini 3.1 Pro Preview",
   "model_id": "google/gemini-3-1-pro-preview",
   "provider": "google",
   "provider_display": "Google DeepMind",
   "score": 51.4,
   "source": "anthropic-system-card-mythos, domain-evals"
  }
 ],
 "page": {
  "aliases": [
   "HLE with search",
   "HLE tool-use"
  ],
  "category": "reasoning",
  "contamination": {
   "note": "Higher practical risk than the no-tools condition, since a model with live search access can retrieve an answer rather than derive it. Anthropic's November 2025 system card describes manually reviewing and re-grading transcripts that showed evidence of this for its own search-enabled runs.\n",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "CC BY 4.0",
   "modalities": [
    "text",
    "image"
   ],
   "public_test_set": true,
   "size": 2500,
   "size_note": "Same question set as hle; see that page for dataset detail.",
   "splits": "single public test split, plus a private held-out set not released",
   "url": "https://huggingface.co/datasets/cais/hle"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "sonnet-5 agent, batch 1b, slice L",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "No dedicated harness task name confirmed for the tool-augmented condition specifically; see hle for the base task."
  },
  "id": "hle_tools",
  "last_updated": "2025-11",
  "leaderboard_url": "https://agi.safe.ai/dashboard",
  "lineage": {
   "family": "hle",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "This id captures Humanity's Last Exam scores produced while the model had tool access, rather than answering closed-book. The specific tool setting varies by reporting source and should be read from that source rather than assumed: Anthropic's Claude Opus 4.5 System Card (November 2025) defines its \"with search\" condition as web search, web fetch and code execution, run without extended thinking, graded by a separate model (Claude Sonnet 4.5) and explicitly decontaminated by flagging transcripts that visited known answer-sheet domains or otherwise showed signs of retrieving rather than deriving an answer. Where a source does not document its tool configuration, that configuration is not established here.\n",
  "metric": {
   "baseline_note": "Directly comparable to hle's accuracy metric on the same question set; not comparable to hle's calibration-error figures, which are not consistently reported for tool-augmented runs.\n",
   "direction": "higher_is_better",
   "max_score": 100,
   "name": "accuracy",
   "unit": "%"
  },
  "name": "Humanity's Last Exam (with tools)",
  "page_kind": "subset",
  "paper": {
   "arxiv": "2501.14249",
   "title": "Humanity's Last Exam",
   "url": "https://arxiv.org/abs/2501.14249",
   "year": 2025
  },
  "publisher": {
   "authors": [],
   "org": "Center for AI Safety and Scale AI",
   "url": "https://lastexam.ai"
  },
  "released": "2025-01",
  "repo_url": "https://github.com/centerforaisafety/hle",
  "saturation": {
   "as_of": "2025-11",
   "note": "Anthropic's Claude Opus 4.5 System Card (Nov 2025) charted six models with search enabled: Claude Opus 4.1 (22.7%), Claude Sonnet 4.5 (28.4%), GPT-5 (35.2%), GPT-5 Pro (42.0%), Claude Opus 4.5 (43.2%) and Gemini 3 (45.8%), each roughly 8-13 points above that same model's own no-tools score in the same chart.\n",
   "status": "open",
   "top_score": 45.8
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "System Card: Claude Opus 4.5 (Anthropic, November 2025), Section 2.16, Figure 2.16.A",
    "url": "https://www.anthropic.com/claude-opus-4-5-system-card"
   },
   {
    "accessed": "2026-09-08",
    "title": "Humanity's Last Exam (arXiv:2501.14249)",
    "url": "https://arxiv.org/abs/2501.14249"
   }
  ],
  "status": "active",
  "subcategory": "frontier academic knowledge Q&A, tool-augmented",
  "summary": "The same 2,500 HLE questions scored when a model can search, fetch web pages and run code, instead of answering from its own knowledge alone.",
  "tags": [
   "reasoning",
   "multi-modal",
   "tool-use",
   "agentic"
  ],
  "task_format": "Same question set and answer format as `hle`, but the model may call tools (web search, web fetch, code execution, or similar, per the reporting source) before producing its final answer.\n"
 }
}