EgyHellaSwag

Egyptian Arabic four-way sentence completion translated from HellaSwag; lm-eval scores the 10,042-item validation split.

Also known as: EgyHellaswag, UBC-NLP/EgyHellaSwag

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryEgyptian Arabic commonsense sentence completion (translated HellaSwag)
Page statusactive
Metricaccuracy (acc); also acc_norm
Directionhigher_is_better
Unit%
Dataset size10052
Dataset licenceMIT
PublisherUBC-NLP (University of British Columbia)

What it measures

EgyHellaSwag is a machine-translated Egyptian Arabic (Masri / ISO arz) version of HellaSwag. Each item gives an activity label and a context sentence plus four endings. The model must pick the plausible continuation. It tests dialectal commonsense sentence completion, not Egyptian cultural knowledge written from scratch. Items keep original HellaSwag source_id values (ActivityNet and WikiHow).

Task format

Four-way multiple choice. lm-eval task egyhellaswag concatenates activity_label and ctx as the query, uses endings as choices, and scores the integer label. Metrics: acc and acc_norm. Training split is 10 rows; scoring uses validation.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub