A tiny, 68-item BIG-bench task modelled on GMAT-style data-sufficiency questions, testing whether a model can tell which of two statements are necessary, sufficient, or redundant to answer a question.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench data-sufficiency task: judging whether given statements are necessary to answer a question (68 items) |
| Page status | superseded |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 68 |
| Dataset licence | Apache-2.0, inherited from the BIG-bench repository as a whole |
Evaluating Information Essentiality poses a question, sometimes with brief context, followed by two supporting statements, and asks the model to judge which combination of those statements is sufficient to answer the question: statement 1 alone, statement 2 alone, either alone, both together, or neither. Unlike a typical exam question that hands the model exactly the information it needs, this task requires the model to first decide whether the information it has been given is enough at all -- a skill the authors argue matters for real-world settings where relevant information is often incomplete or mixed with irrelevant detail, closer to a GMAT-style "data sufficiency" question than a standard reading-comprehension item.
Five-option multiple choice, scored with BIG-bench's `multiple_choice_grade` metric; every one of the 68 items uses the same five-way answer structure (64 items share one canonical wording, 4 substitute "the question can be answered without either statement" for the "neither is sufficient" option).
No model card in ModelSpec reports this benchmark yet.