Humanity's Last Exam (with tools)

The same 2,500 HLE questions scored when a model can search, fetch web pages and run code, instead of answering from its own knowledge alone.

Also known as: HLE with search, HLE tool-use

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryfrontier academic knowledge Q&A, tool-augmented
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size2500
Dataset licenceCC BY 4.0
PublisherCenter for AI Safety and Scale AI

What it measures

This id captures Humanity's Last Exam scores produced while the model had tool access, rather than answering closed-book. The specific tool setting varies by reporting source and should be read from that source rather than assumed: Anthropic's Claude Opus 4.5 System Card (November 2025) defines its "with search" condition as web search, web fetch and code execution, run without extended thinking, graded by a separate model (Claude Sonnet 4.5) and explicitly decontaminated by flagging transcripts that visited known answer-sheet domains or otherwise showed signs of retrieving rather than deriving an answer. Where a source does not document its tool configuration, that configuration is not established here.

Task format

Same question set and answer format as `hle`, but the model may call tools (web search, web fetch, code execution, or similar, per the reporting source) before producing its final answer.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic64.72026-04
Claude Opus 4.6Anthropic53.12026-04
GLM 5.1Z.ai (Zhipu AI)52.32026-04
GPT-5.4OpenAI52.12026-04
Gemini 3.1 Pro PreviewGoogle DeepMind51.42026-04

Data

This page as JSON · Edit on GitHub