Word Problems on Sets and Graphs

A BIG-bench free-response task that checks whether a model tracks set membership and graph paths across three small synthetic puzzle types.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorylogical reasoning
Page statusactive
Metrickeyword-match score
Directionhigher_is_better
Unitpoints
Dataset size3000
Dataset licenceApache-2.0

What it measures

BIG-bench's word_problems_on_sets_and_graphs task tests whether a model can track set membership, graph reachability, and set operations across three synthetic word-problem styles, without relying on surface word co-occurrence.

Task format

Free-text generation. The model reads a short narrative (e.g. fruits added to and removed from a basket) and must produce the final answer as free text; responses are keyword-checked rather than graded on exact wording or grammar.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub