1,500 BIG-bench questions that ask which object sits at a given position after a minimal set of ordering clues; also three 250-item BBH tasks.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench (and BBH) object-ordering deduction from a minimal clue set |
| Page status | active |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1500 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
logical_deduction describes three, five or seven similar objects in a natural order (books on a shelf, golfers in a ranking, fruit by price) and gives a minimal set of clues such as "X is left of Y" or "X is third." The model must assign higher probability to the true statement about which object occupies a queried position than to the false statements for the other objects. Authors James Simon and Chandan Singh generated the puzzles so each has exactly one solution and no spare clue. The task is meant to need multi-step deduction rather than pattern match.
Multiple-choice over N statements for an N-object puzzle, scored with `multiple_choice_grade`. BIG-bench ships three subtasks (three_objects, five_objects, seven_objects). A canary GUID is embedded. BIG-bench Hard keeps three separate 250-example files under the same names.
No model card in ModelSpec reports this benchmark yet.