Okapi multilingual HellaSwag

lm-eval group of GPT-translated HellaSwag val sets in 30 languages; four-way continuation, scored acc and acc_norm.

Also known as: hellaswag_multilingual, m_hellaswag, alexandrainst/m_hellaswag

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorymachine-translated four-way commonsense sentence continuation
Page statusactive
Metricaccuracy (acc and length-normalised acc_norm)
Directionhigher_is_better
Unit%
Dataset size275384
Dataset licenceCC-BY-NC-4.0
PublisherUniversity of Oregon NLP (Okapi translations); Alexandra Institute (Hub dump); EleutherAI (lm-eval group)

What it measures

okapi_hellaswag_multilingual is EleutherAI's lm-evaluation-harness group over machine-translated HellaSwag. The model reads an activity label plus a short scene and must pick which of four endings is the everyday next step. Items come from ActivityNet captions and WikiHow, same as English [HellaSwag](hellaswag.md). University of Oregon translated the English set with ChatGPT (GPT-3.5-turbo) for the Okapi paper; Alexandra Institute hosts the Hub dump and added Icelandic and Norwegian. Text only. This is not native-writer Darija or Egyptian Arabic HellaSwag.

Task format

Four-way multiple choice. Runnable tasks are hellaswag_{ar,bn,ca,da,de,es,eu,fr, gu,hi,hr,hu,hy,id,it,kn,ml,mr,ne,nl,pt,ro,ru,sk,sr,sv,ta,te,uk,vi}. YAML tag and README group: hellaswag_multilingual. Prompt is activity_label + ": " + ctx_a/ctx_b after WikiHow-bracket cleanup; target is the original label field; metrics acc and acc_norm. Split is val only.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub