Evaluating Information Essentiality

A tiny, 68-item BIG-bench task modelled on GMAT-style data-sufficiency questions, testing whether a model can tell which of two statements are necessary, sufficient, or redundant to answer a question.

Also known as: The Essential, the Excessive, and the Extraneous

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench data-sufficiency task: judging whether given statements are necessary to answer a question (68 items)
Page statussuperseded
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size68
Dataset licenceApache-2.0, inherited from the BIG-bench repository as a whole

What it measures

Evaluating Information Essentiality poses a question, sometimes with brief context, followed by two supporting statements, and asks the model to judge which combination of those statements is sufficient to answer the question: statement 1 alone, statement 2 alone, either alone, both together, or neither. Unlike a typical exam question that hands the model exactly the information it needs, this task requires the model to first decide whether the information it has been given is enough at all -- a skill the authors argue matters for real-world settings where relevant information is often incomplete or mixed with irrelevant detail, closer to a GMAT-style "data sufficiency" question than a standard reading-comprehension item.

Task format

Five-option multiple choice, scored with BIG-bench's `multiple_choice_grade` metric; every one of the 68 items uses the same five-way answer structure (64 items share one canonical wording, 4 substitute "the question can be answered without either statement" for the "neither is sufficient" option).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub