Replaces each of BBH's 23 tasks with a substantially harder variant of the same reasoning skill, calibrated so two strong 2025 reference models both scored under 70%.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | extra-hard multi-step reasoning suite, BBH successor |
| Page status | active |
| Metric | harmonic mean accuracy across the 23 tasks (classic micro-average accuracy recommended for the smaller Mini split) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 4520 |
| Dataset licence | Apache-2.0 (evaluation code); CC BY 4.0 International (other materials) |
| Publisher | Google DeepMind |
BIG-Bench Extra Hard (BBEH) takes the same 23 task categories BIG-Bench Hard (BBH) uses -- logical deduction, causal judgement, object tracking and counting, spatial and temporal reasoning, disambiguation, and several more -- and replaces every one of BBH's individual tasks with a new, harder task designed to probe the same underlying reasoning skill. The authors built each replacement task iteratively, testing candidate items against two Google reference models (Gemini 1.5 Flash and a Gemini "Thinking Experimental" model) and refining until both scored below 70% accuracy, so difficulty is calibrated against contemporary models rather than guessed. The result targets the same reasoning categories as BBH while addressing the saturation BBH itself had started to show against increasingly strong models.
A mix of multiple-choice and free-response prompts across 23 tasks, mirroring BBH's task categories but with harder instances; most are answered directly or via chain-of-thought prompting before a final extracted answer.
No model card in ModelSpec reports this benchmark yet.