WSC273

WSC273 scores pronoun-resolution accuracy on the first 273 items of the Winograd Schema Challenge, using language-model probability rather than fine-tuning.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorycoreference resolution
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unitpercent
Dataset size273
Dataset licencecc-by-4.0

What it measures

WSC273 evaluates commonsense pronoun disambiguation on the first 273 items of the Winograd Schema Challenge, a set of near-identical sentence pairs whose correct pronoun referent flips on world knowledge alone.

Task format

Multiple-choice: given a sentence with an ambiguous pronoun and two candidate referents that differ by one or two words from a paired sentence, the model must pick the correct referent, scored via language-model probability of the completion (partial evaluation) rather than explicit answer selection.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub