A 44k-problem, adversarially filtered successor to the Winograd Schema Challenge, testing commonsense pronoun resolution at scale.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | commonsense coreference resolution |
| Page status | unknown |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 44000 |
| Dataset licence | CC-BY |
| Publisher | Allen Institute for AI (AI2) |
WinoGrande gives a model a short sentence with a blank that must be filled with one of two candidate nouns or phrases, where picking correctly requires commonsense reasoning about the situation rather than grammar or word association. It scales up the original, hand-crafted 273-problem Winograd Schema Challenge to 44,000 problems using crowdsourcing plus an adversarial filtering algorithm (AfLite) that removes items solvable by superficial statistical shortcuts, specifically to stop models from succeeding via spurious dataset bias rather than genuine commonsense understanding.
Binary fill-in-the-blank choice between two candidate answers for a short sentence.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 85.2 | 2026-04 |
| Yi 1.5 34B | 01.AI | 84.1 | 2026-04 |
| Yi 1.5 34B 32K | 01.AI | 84.0 | 2024-07 |
| Nous Hermes 2 Mixtral 8x7B DPO | Nous Research | 83.1 | 2024-07 |
| Nous Hermes 2 Yi 34B | Nous Research | 83.0 | 2024-07 |
| Yi 34B | 01.AI | 83.0 | 2024-07 |
| Meta Llama 3 70B | Meta | 82.9 | 2024-07 |
| Meta Llama 3 70B Instruct | Meta | 82.9 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 82.9 | 2026-04 |
| Yi 34B 200K | 01.AI | 82.9 | 2024-07 |
| Nous Hermes 2 SOLAR 10.7B | Nous Research | 82.8 | 2024-07 |
| Mixtral 8x7B v0.1 | Mistral AI | 81.7 | 2026-04 |
| Yi 1.5 34B Chat | 01.AI | 81.6 | 2026-04 |
| Yi 1.5 34B Chat 16K | 01.AI | 81.6 | 2024-07 |
| Mixtral 8x7B Instruct v0.1 | Mistral AI | 81.1 | 2026-04 |
| Llama 2 70B chat hf | Meta | 80.5 | 2024-07 |
| Yi 34B Chat | 01.AI | 80.1 | 2024-07 |
| gemma 7B it | Google DeepMind | 78.5 | 2024-07 |
| Meta Llama 3 8B | Meta | 78.5 | 2024-07 |
| Meta Llama 3 8B | Nous Research | 78.5 | 2024-07 |
| Mistral 7B v0.3 | Mistral AI | 78.5 | 2024-07 |
| mistral 7B v0.3 bnb 4bit | Unsloth | 78.5 | 2024-07 |
| falcon 11B | TII | 78.3 | 2024-07 |
| Yi 1.5 9B | 01.AI | 78.0 | 2024-07 |
| Yi 9B | 01.AI | 77.5 | 2024-07 |
| Mistral 7B Instruct v0.2 | Mistral AI | 77.2 | 2024-07 |
| Yi 1.5 9B Chat | 01.AI | 77.2 | 2024-07 |
| Yi 1.5 9B 32K | 01.AI | 76.9 | 2024-07 |
| Hermes 2 Theta Llama 3 8B | Nous Research | 76.6 | 2024-07 |
| Hermes 2 Pro Llama 3 8B | Nous Research | 76.4 | 2024-07 |
| Yi 1.5 6B | 01.AI | 75.5 | 2024-07 |
| Yi 1.5 9B Chat 16K | 01.AI | 75.0 | 2024-07 |
| deepseek llm 7B base | DeepSeek | 74.9 | 2024-07 |
| deepseek llm 7B chat | DeepSeek | 74.9 | 2024-07 |
| Llama 2 13B chat hf | Meta | 74.5 | 2024-07 |
| Meta Llama 3 8B Instruct | Meta | 74.5 | 2024-07 |
| Meta Llama 3 8B Instruct | Nous Research | 74.5 | 2024-07 |
| Yi 6B | 01.AI | 74.2 | 2024-07 |
| Yi 6B Chat | 01.AI | 74.2 | 2024-07 |
| Nous Hermes llama 2 7B | Nous Research | 74.0 | 2024-07 |
| Mistral 7B Instruct v0.1 | Mistral AI | 73.7 | 2024-07 |
| Yi 1.5 6B Chat | 01.AI | 73.6 | 2024-07 |
| phi 2 | Microsoft | 73.5 | 2024-07 |
| CodeLlama 34B Instruct hf | Meta | 73.4 | 2026-04 |
| Phi 3 mini 128K instruct | Microsoft | 72.8 | 2024-07 |
| Phi 3 mini 4K instruct | Microsoft | 72.4 | 2024-07 |
| Llama 2 7B chat hf | Meta | 71.7 | 2024-07 |
| Llama 2 7B chat hf | Nous Research | 71.7 | 2024-07 |
| Baichuan 7B | Baichuan | 66.8 | 2024-07 |
| gemma 2B | Google DeepMind | 66.3 | 2024-07 |
| CodeLlama 7B hf | Meta | 64.9 | 2024-07 |
| CodeLlama 7B Instruct hf | Meta | 64.9 | 2024-07 |
| Qwen2 1.5B Instruct | Alibaba / Qwen Team | 64.9 | 2024-07 |
| OLMo 1B hf | Allen AI | 61.1 | 2024-07 |
| gemma 2B it | Google DeepMind | 60.9 | 2024-07 |
| deepseek coder 6.7B base | DeepSeek | 58.1 | 2024-07 |
| deepseek coder 6.7B instruct | DeepSeek | 56.8 | 2024-07 |
| Qwen2 0.5B Instruct | Alibaba / Qwen Team | 55.8 | 2024-07 |
| deepseek coder 1.3B base | DeepSeek | 52.4 | 2024-07 |
| deepseek coder 1.3B instruct | DeepSeek | 52.4 | 2024-07 |