WSC273 scores pronoun-resolution accuracy on the first 273 items of the Winograd Schema Challenge, using language-model probability rather than fine-tuning.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | coreference resolution |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 273 |
| Dataset licence | cc-by-4.0 |
WSC273 evaluates commonsense pronoun disambiguation on the first 273 items of the Winograd Schema Challenge, a set of near-identical sentence pairs whose correct pronoun referent flips on world knowledge alone.
Multiple-choice: given a sentence with an ambiguous pronoun and two candidate referents that differ by one or two words from a paired sentence, the model must pick the correct referent, scored via language-model probability of the completion (partial evaluation) rather than explicit answer selection.
No model card in ModelSpec reports this benchmark yet.