HellaSwag

Four-way multiple-choice test of commonsense sentence continuation, built by adversarial filtering to defeat models of its era.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorycommonsense natural language inference
Page statussaturated
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size59950
Dataset licenceMIT
PublisherPaul G. Allen School of Computer Science & Engineering, University of Washington

What it measures

HellaSwag tests commonsense inference: given a short description of an everyday situation, a model must pick which of four possible continuations is the most plausible next step. Contexts are drawn from ActivityNet video captions and WikiHow how-to articles, covering ordinary physical activities such as cooking, repairs or sports. The task looks trivial to a person but was built specifically to be hard for the language models of its time, isolating whether a model has a working sense of how everyday situations unfold rather than just fluent language generation.

Task format

Four-way multiple choice; given a short context, the model selects the most plausible of four candidate continuations.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Mixtral 8x22B Instruct v0.1Mistral AI89.12026-04
Mixtral 8x7B Instruct v0.1Mistral AI87.62026-04
Nous Hermes 2 Mixtral 8x7B DPONous Research87.32024-07
Mixtral 8x7B v0.1Mistral AI86.52026-04
Yi 1.5 34B01.AI86.12026-04
Yi 1.5 34B Chat01.AI86.02026-04
Yi 1.5 34B Chat 16K01.AI86.02024-07
Meta Llama 3 70BMeta85.72024-07
Meta Llama 3 70B InstructMeta85.72026-04
Meta Llama 3 70B InstructNous Research85.72026-04
Nous Hermes 2 Yi 34BNous Research85.52024-07
Yi 1.5 34B 32K01.AI85.52024-07
falcon 40BTII85.32024-07
Mistral 7B Instruct v0.2Mistral AI84.92024-07
Nous Hermes 2 SOLAR 10.7BNous Research84.92024-07
Yi 34B Chat01.AI84.22024-07
Hermes 2 Pro Llama 3 8BNous Research83.22024-07
Mistral 7B v0.3Mistral AI83.02024-07
mistral 7B v0.3 bnb 4bitUnsloth83.02024-07
Hermes 2 Theta Llama 3 8BNous Research82.62024-07
gemma 7B itGoogle DeepMind82.52024-07
Meta Llama 3 8BMeta82.22024-07
Meta Llama 3 8BNous Research82.22024-07
Yi 34B 200K01.AI82.12024-07
Yi 1.5 9B Chat01.AI80.92024-07
Yi 1.5 9B Chat 16K01.AI80.92024-07
Phi 3 mini 4K instructMicrosoft80.62024-07
Yi 1.5 9B01.AI80.42024-07
Phi 3 mini 128K instructMicrosoft80.12024-07
Yi 1.5 9B 32K01.AI79.82024-07
deepseek llm 7B baseDeepSeek79.42024-07
deepseek llm 7B chatDeepSeek79.42024-07
Yi 1.5 6B Chat01.AI78.92024-07
Yi 9B01.AI78.82024-07
Meta Llama 3 8B InstructMeta78.62024-07
Meta Llama 3 8B InstructNous Research78.62024-07
Yi 1.5 6B01.AI78.02024-07
Yi 6B01.AI76.62024-07
Yi 6B Chat01.AI76.62024-07
phi 2Microsoft74.92024-07
gemma 2BGoogle DeepMind71.82024-07
Qwen2 1.5B InstructAlibaba / Qwen Team66.72024-07
OLMo 1B hfAllen AI63.62024-07
gemma 2B itGoogle DeepMind62.72024-07
chatglm2 6BZhipu AI59.02024-07
deepseek coder 6.7B instructDeepSeek55.12024-07
deepseek coder 6.7B baseDeepSeek53.52024-07
Qwen2 0.5B InstructAlibaba / Qwen Team48.72024-07
deepseek coder 1.3B baseDeepSeek39.92024-07
deepseek coder 1.3B instructDeepSeek39.92024-07

Data

This page as JSON · Edit on GitHub