WinoWhy tests whether a model can pick the correct justification for a Winograd Schema Challenge answer, not just the answer itself.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | reasoning |
| Page status | active |
| Metric | multiple choice grade |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 2862 |
| Dataset licence | Apache-2.0 |
BIG-bench's winowhy task asks a model to distinguish the correct commonsense reason for a Winograd Schema Challenge pronoun resolution from plausible-looking wrong reasons.
Multiple-choice; given a WSC sentence pair and its correct pronoun resolution, the model must pick the correct justification for that resolution from several candidate reasons.
No model card in ModelSpec reports this benchmark yet.