2,500 expert-written, closed-ended questions spanning dozens of academic subjects, built by CAIS and Scale AI to replace saturated benchmarks like MMLU.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | frontier academic knowledge Q&A |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2500 |
| Dataset licence | CC BY 4.0 |
| Publisher | Center for AI Safety and Scale AI |
Humanity's Last Exam (HLE) tests whether a model can answer closed-ended academic questions at the frontier of human expertise, across dozens of subjects including mathematics, the humanities and the natural sciences. Each question was written by a subject-matter expert and is deliberately constructed to have a single, unambiguous, verifiable answer that cannot be produced quickly by searching the internet. About 90% of questions are text-only; the remainder pair text with a reference image, so the benchmark is described by its authors as multi-modal rather than vision-first. The benchmark also scores calibration: whether a model's stated confidence matches how often it is actually correct.
Single-turn question, either exact-match (a short string or number the model must produce, roughly 80% of items) or multiple-choice with five or more options (the remainder). About 10% of questions include a reference image alongside the text. Models are also asked to state a numeric confidence (0-100%) alongside their answer, which feeds the calibration metric.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 56.8 | 2026-04 |
| Muse Spark | Meta | 50.2 | 2026-04 |
| Gemini 3.1 Pro Preview | Google DeepMind | 44.4 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 40.0 | 2026-04 |
| GPT-5.4 | OpenAI | 39.8 | 2026-04 |
| GLM 5.1 | Z.ai (Zhipu AI) | 31.0 | 2026-04 |