An 11-task, four-way suite of held-out test slices for Moroccan Darija covering sentiment analysis, summarization, translation and transliteration, built for the Atlas-Chat model release.
unassessed
| Category | composite |
|---|---|
| Subcategory | multi-task Moroccan Darija NLP suite: sentiment analysis, summarization, translation and transliteration |
| Page status | active |
| Metric | per-task-group metric (accuracy for sentiment; ROUGE/BERTScore/chrF for summarization; BLEU/chrF/TER/BERTScore for translation; BERTScore for transliteration) |
| Direction | higher_is_better |
| Unit | mixed (%, points) |
| Dataset size | 31476 |
| Dataset licence | ODC-BY (top-level mirror licence); the dataset card explicitly warns this is a mixture with component-level licences that differ, some non-commercial -- see Dataset and licence. |
| Publisher | MBZUAI-Paris, in a multi-institution collaboration with EMINES-UM6P, LINAGORA, KTH, AtlasIA and École Polytechnique |
DarijaBench bundles held-out test portions of pre-existing Darija datasets into four task groups: sentiment analysis (5 source datasets), summarization (1), translation (4, across six Darija-MSA/ English/French directions) and transliteration between Darija written in Arabic script and Arabizi written in Latin script (1). It does not test one skill; it is a fixed collection point for Darija-specific NLP evaluation, assembled because, as the authors state, standardized Darija benchmarks barely existed beforehand. Most component test portions are a 10% held-out split of their respective source dataset.
Format varies by group: sentiment is multiple-class classification (accuracy); summarization is free-text generation from a Darija article (ROUGE-1/2/L/Lsum, a BERTScore-style metric, and chrF); translation is free-text generation across six directions (BLEU, chrF, TER and a BERTScore-style metric); transliteration is free-text script conversion (the same BERTScore-style metric). All are zero-shot, single-turn generation tasks with no worked examples in the released harness configs.
No model card in ModelSpec reports this benchmark yet.