WinoGrande

A 44k-problem, adversarially filtered successor to the Winograd Schema Challenge, testing commonsense pronoun resolution at scale.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorycommonsense coreference resolution
Page statusunknown
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size44000
Dataset licenceCC-BY
PublisherAllen Institute for AI (AI2)

What it measures

WinoGrande gives a model a short sentence with a blank that must be filled with one of two candidate nouns or phrases, where picking correctly requires commonsense reasoning about the situation rather than grammar or word association. It scales up the original, hand-crafted 273-problem Winograd Schema Challenge to 44,000 problems using crowdsourcing plus an adversarial filtering algorithm (AfLite) that removes items solvable by superficial statistical shortcuts, specifically to stop models from succeeding via spurious dataset bias rather than genuine commonsense understanding.

Task format

Binary fill-in-the-blank choice between two candidate answers for a short sentence.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Mixtral 8x22B Instruct v0.1Mistral AI85.22026-04
Yi 1.5 34B01.AI84.12026-04
Yi 1.5 34B 32K01.AI84.02024-07
Nous Hermes 2 Mixtral 8x7B DPONous Research83.12024-07
Nous Hermes 2 Yi 34BNous Research83.02024-07
Yi 34B01.AI83.02024-07
Meta Llama 3 70BMeta82.92024-07
Meta Llama 3 70B InstructMeta82.92026-04
Meta Llama 3 70B InstructNous Research82.92026-04
Yi 34B 200K01.AI82.92024-07
Nous Hermes 2 SOLAR 10.7BNous Research82.82024-07
Mixtral 8x7B v0.1Mistral AI81.72026-04
Yi 1.5 34B Chat01.AI81.62026-04
Yi 1.5 34B Chat 16K01.AI81.62024-07
Mixtral 8x7B Instruct v0.1Mistral AI81.12026-04
Llama 2 70B chat hfMeta80.52024-07
Yi 34B Chat01.AI80.12024-07
gemma 7B itGoogle DeepMind78.52024-07
Meta Llama 3 8BMeta78.52024-07
Meta Llama 3 8BNous Research78.52024-07
Mistral 7B v0.3Mistral AI78.52024-07
mistral 7B v0.3 bnb 4bitUnsloth78.52024-07
falcon 11BTII78.32024-07
Yi 1.5 9B01.AI78.02024-07
Yi 9B01.AI77.52024-07
Mistral 7B Instruct v0.2Mistral AI77.22024-07
Yi 1.5 9B Chat01.AI77.22024-07
Yi 1.5 9B 32K01.AI76.92024-07
Hermes 2 Theta Llama 3 8BNous Research76.62024-07
Hermes 2 Pro Llama 3 8BNous Research76.42024-07
Yi 1.5 6B01.AI75.52024-07
Yi 1.5 9B Chat 16K01.AI75.02024-07
deepseek llm 7B baseDeepSeek74.92024-07
deepseek llm 7B chatDeepSeek74.92024-07
Llama 2 13B chat hfMeta74.52024-07
Meta Llama 3 8B InstructMeta74.52024-07
Meta Llama 3 8B InstructNous Research74.52024-07
Yi 6B01.AI74.22024-07
Yi 6B Chat01.AI74.22024-07
Nous Hermes llama 2 7BNous Research74.02024-07
Mistral 7B Instruct v0.1Mistral AI73.72024-07
Yi 1.5 6B Chat01.AI73.62024-07
phi 2Microsoft73.52024-07
CodeLlama 34B Instruct hfMeta73.42026-04
Phi 3 mini 128K instructMicrosoft72.82024-07
Phi 3 mini 4K instructMicrosoft72.42024-07
Llama 2 7B chat hfMeta71.72024-07
Llama 2 7B chat hfNous Research71.72024-07
Baichuan 7BBaichuan66.82024-07
gemma 2BGoogle DeepMind66.32024-07
CodeLlama 7B hfMeta64.92024-07
CodeLlama 7B Instruct hfMeta64.92024-07
Qwen2 1.5B InstructAlibaba / Qwen Team64.92024-07
OLMo 1B hfAllen AI61.12024-07
gemma 2B itGoogle DeepMind60.92024-07
deepseek coder 6.7B baseDeepSeek58.12024-07
deepseek coder 6.7B instructDeepSeek56.82024-07
Qwen2 0.5B InstructAlibaba / Qwen Team55.82024-07
deepseek coder 1.3B baseDeepSeek52.42024-07
deepseek coder 1.3B instructDeepSeek52.42024-07

Data

This page as JSON · Edit on GitHub