Darija translation of HellaSwag: pick the plausible ending among four Moroccan Arabic continuations of a short scene.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | Moroccan Darija four-way commonsense sentence continuation |
| Page status | active |
| Metric | accuracy (acc and length-normalised acc_norm) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 20055 |
| Dataset licence | MIT |
| Publisher | MBZUAI-Paris, with EMINES-UM6P, LINAGORA, KTH, AtlasIA and École Polytechnique |
DarijaHellaSwag is HellaSwag rewritten in Moroccan Darija. The model reads a short scene (activity label plus context) and must choose which of four endings is the everyday continuation. The Atlas-Chat paper and the Hugging Face card both say Claude 3.5 Sonnet translated the English HellaSwag validation set, with native-speaker review. The Hub snapshot also ships a 10,003-row test split matching original HellaSwag test size, plus a 10-row train split. lm-evaluation-harness scores the validation split.
Four-way multiple choice. lm-eval task darijahellaswag builds query as activity_label + ": " + ctx, choices from endings, metrics acc and acc_norm. training_split train (10 rows), validation_split validation, test_split null.
No model card in ModelSpec reports this benchmark yet.