Four-way multiple-choice test of commonsense sentence continuation, built by adversarial filtering to defeat models of its era.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | commonsense natural language inference |
| Page status | saturated |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 59950 |
| Dataset licence | MIT |
| Publisher | Paul G. Allen School of Computer Science & Engineering, University of Washington |
HellaSwag tests commonsense inference: given a short description of an everyday situation, a model must pick which of four possible continuations is the most plausible next step. Contexts are drawn from ActivityNet video captions and WikiHow how-to articles, covering ordinary physical activities such as cooking, repairs or sports. The task looks trivial to a person but was built specifically to be hard for the language models of its time, isolating whether a model has a working sense of how everyday situations unfold rather than just fluent language generation.
Four-way multiple choice; given a short context, the model selects the most plausible of four candidate continuations.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 89.1 | 2026-04 |
| Mixtral 8x7B Instruct v0.1 | Mistral AI | 87.6 | 2026-04 |
| Nous Hermes 2 Mixtral 8x7B DPO | Nous Research | 87.3 | 2024-07 |
| Mixtral 8x7B v0.1 | Mistral AI | 86.5 | 2026-04 |
| Yi 1.5 34B | 01.AI | 86.1 | 2026-04 |
| Yi 1.5 34B Chat | 01.AI | 86.0 | 2026-04 |
| Yi 1.5 34B Chat 16K | 01.AI | 86.0 | 2024-07 |
| Meta Llama 3 70B | Meta | 85.7 | 2024-07 |
| Meta Llama 3 70B Instruct | Meta | 85.7 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 85.7 | 2026-04 |
| Nous Hermes 2 Yi 34B | Nous Research | 85.5 | 2024-07 |
| Yi 1.5 34B 32K | 01.AI | 85.5 | 2024-07 |
| falcon 40B | TII | 85.3 | 2024-07 |
| Mistral 7B Instruct v0.2 | Mistral AI | 84.9 | 2024-07 |
| Nous Hermes 2 SOLAR 10.7B | Nous Research | 84.9 | 2024-07 |
| Yi 34B Chat | 01.AI | 84.2 | 2024-07 |
| Hermes 2 Pro Llama 3 8B | Nous Research | 83.2 | 2024-07 |
| Mistral 7B v0.3 | Mistral AI | 83.0 | 2024-07 |
| mistral 7B v0.3 bnb 4bit | Unsloth | 83.0 | 2024-07 |
| Hermes 2 Theta Llama 3 8B | Nous Research | 82.6 | 2024-07 |
| gemma 7B it | Google DeepMind | 82.5 | 2024-07 |
| Meta Llama 3 8B | Meta | 82.2 | 2024-07 |
| Meta Llama 3 8B | Nous Research | 82.2 | 2024-07 |
| Yi 34B 200K | 01.AI | 82.1 | 2024-07 |
| Yi 1.5 9B Chat | 01.AI | 80.9 | 2024-07 |
| Yi 1.5 9B Chat 16K | 01.AI | 80.9 | 2024-07 |
| Phi 3 mini 4K instruct | Microsoft | 80.6 | 2024-07 |
| Yi 1.5 9B | 01.AI | 80.4 | 2024-07 |
| Phi 3 mini 128K instruct | Microsoft | 80.1 | 2024-07 |
| Yi 1.5 9B 32K | 01.AI | 79.8 | 2024-07 |
| deepseek llm 7B base | DeepSeek | 79.4 | 2024-07 |
| deepseek llm 7B chat | DeepSeek | 79.4 | 2024-07 |
| Yi 1.5 6B Chat | 01.AI | 78.9 | 2024-07 |
| Yi 9B | 01.AI | 78.8 | 2024-07 |
| Meta Llama 3 8B Instruct | Meta | 78.6 | 2024-07 |
| Meta Llama 3 8B Instruct | Nous Research | 78.6 | 2024-07 |
| Yi 1.5 6B | 01.AI | 78.0 | 2024-07 |
| Yi 6B | 01.AI | 76.6 | 2024-07 |
| Yi 6B Chat | 01.AI | 76.6 | 2024-07 |
| phi 2 | Microsoft | 74.9 | 2024-07 |
| gemma 2B | Google DeepMind | 71.8 | 2024-07 |
| Qwen2 1.5B Instruct | Alibaba / Qwen Team | 66.7 | 2024-07 |
| OLMo 1B hf | Allen AI | 63.6 | 2024-07 |
| gemma 2B it | Google DeepMind | 62.7 | 2024-07 |
| chatglm2 6B | Zhipu AI | 59.0 | 2024-07 |
| deepseek coder 6.7B instruct | DeepSeek | 55.1 | 2024-07 |
| deepseek coder 6.7B base | DeepSeek | 53.5 | 2024-07 |
| Qwen2 0.5B Instruct | Alibaba / Qwen Team | 48.7 | 2024-07 |
| deepseek coder 1.3B base | DeepSeek | 39.9 | 2024-07 |
| deepseek coder 1.3B instruct | DeepSeek | 39.9 | 2024-07 |