Humanity's Last Exam

2,500 expert-written, closed-ended questions spanning dozens of academic subjects, built by CAIS and Scale AI to replace saturated benchmarks like MMLU.

Also known as: HLE

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryfrontier academic knowledge Q&A
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size2500
Dataset licenceCC BY 4.0
PublisherCenter for AI Safety and Scale AI

What it measures

Humanity's Last Exam (HLE) tests whether a model can answer closed-ended academic questions at the frontier of human expertise, across dozens of subjects including mathematics, the humanities and the natural sciences. Each question was written by a subject-matter expert and is deliberately constructed to have a single, unambiguous, verifiable answer that cannot be produced quickly by searching the internet. About 90% of questions are text-only; the remainder pair text with a reference image, so the benchmark is described by its authors as multi-modal rather than vision-first. The benchmark also scores calibration: whether a model's stated confidence matches how often it is actually correct.

Task format

Single-turn question, either exact-match (a short string or number the model must produce, roughly 80% of items) or multiple-choice with five or more options (the remainder). About 10% of questions include a reference image alongside the text. Models are also asked to state a numeric confidence (0-100%) alongside their answer, which feeds the calibration metric.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic56.82026-04
Muse SparkMeta50.22026-04
Gemini 3.1 Pro PreviewGoogle DeepMind44.42026-04
Claude Opus 4.6Anthropic40.02026-04
GPT-5.4OpenAI39.82026-04
GLM 5.1Z.ai (Zhipu AI)31.02026-04

Data

This page as JSON · Edit on GitHub