Twenty Questions

Two model instances play Twenty Questions, communicating a hidden concept through yes-or-no answers.

Also known as: BIG-bench Twenty Questions

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryself-play concept identification
Page statusactive
Metricnegative number of conversational rounds to the correct guess
Directionhigher_is_better
Unitrounds
PublisherBIG-bench collaboration

What it measures

Alice receives a hidden concept and answers yes or no. Bob asks questions and must identify it. The task tests constrained answering, targeted question selection, persistence and self-play.

Task format

Interactive two-agent text game with one question per turn and a maximum of 100 questions.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub