WinoWhy

WinoWhy tests whether a model can pick the correct justification for a Winograd Schema Challenge answer, not just the answer itself.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategoryreasoning
Page statusactive
Metricmultiple choice grade
Directionhigher_is_better
Unitpercent
Dataset size2862
Dataset licenceApache-2.0

What it measures

BIG-bench's winowhy task asks a model to distinguish the correct commonsense reason for a Winograd Schema Challenge pronoun resolution from plausible-looking wrong reasons.

Task format

Multiple-choice; given a WSC sentence pair and its correct pronoun resolution, the model must pick the correct justification for that resolution from several candidate reasons.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub