BIG-Bench Extra Hard (BBEH)

Replaces each of BBH's 23 tasks with a substantially harder variant of the same reasoning skill, calibrated so two strong 2025 reference models both scored under 70%.

Also known as: BBEH

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryextra-hard multi-step reasoning suite, BBH successor
Page statusactive
Metricharmonic mean accuracy across the 23 tasks (classic micro-average accuracy recommended for the smaller Mini split)
Directionhigher_is_better
Unit%
Dataset size4520
Dataset licenceApache-2.0 (evaluation code); CC BY 4.0 International (other materials)
PublisherGoogle DeepMind

What it measures

BIG-Bench Extra Hard (BBEH) takes the same 23 task categories BIG-Bench Hard (BBH) uses -- logical deduction, causal judgement, object tracking and counting, spatial and temporal reasoning, disambiguation, and several more -- and replaces every one of BBH's individual tasks with a new, harder task designed to probe the same underlying reasoning skill. The authors built each replacement task iteratively, testing candidate items against two Google reference models (Gemini 1.5 Flash and a Gemini "Thinking Experimental" model) and refining until both scored below 70% accuracy, so difficulty is calibrated against contemporary models rather than guessed. The result targets the same reasoning categories as BBH while addressing the saturation BBH itself had started to show against increasingly strong models.

Task format

A mix of multiple-choice and free-response prompts across 23 tasks, mirroring BBH's task categories but with harder instances; most are answered directly or via chain-of-thought prompting before a final extracted answer.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub