A BIG-bench free-response task that checks whether a model tracks set membership and graph paths across three small synthetic puzzle types.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | logical reasoning |
| Page status | active |
| Metric | keyword-match score |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 3000 |
| Dataset licence | Apache-2.0 |
BIG-bench's word_problems_on_sets_and_graphs task tests whether a model can track set membership, graph reachability, and set operations across three synthetic word-problem styles, without relying on surface word co-occurrence.
Free-text generation. The model reads a short narrative (e.g. fruits added to and removed from a basket) and must produce the final answer as free text; responses are keyword-checked rather than graded on exact wording or grammar.
No model card in ModelSpec reports this benchmark yet.