The same 2,500 HLE questions scored when a model can search, fetch web pages and run code, instead of answering from its own knowledge alone.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | frontier academic knowledge Q&A, tool-augmented |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2500 |
| Dataset licence | CC BY 4.0 |
| Publisher | Center for AI Safety and Scale AI |
This id captures Humanity's Last Exam scores produced while the model had tool access, rather than answering closed-book. The specific tool setting varies by reporting source and should be read from that source rather than assumed: Anthropic's Claude Opus 4.5 System Card (November 2025) defines its "with search" condition as web search, web fetch and code execution, run without extended thinking, graded by a separate model (Claude Sonnet 4.5) and explicitly decontaminated by flagging transcripts that visited known answer-sheet domains or otherwise showed signs of retrieving rather than deriving an answer. Where a source does not document its tool configuration, that configuration is not established here.
Same question set and answer format as `hle`, but the model may call tools (web search, web fetch, code execution, or similar, per the reporting source) before producing its final answer.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 64.7 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 53.1 | 2026-04 |
| GLM 5.1 | Z.ai (Zhipu AI) | 52.3 | 2026-04 |
| GPT-5.4 | OpenAI | 52.1 | 2026-04 |
| Gemini 3.1 Pro Preview | Google DeepMind | 51.4 | 2026-04 |