DarijaBench

An 11-task, four-way suite of held-out test slices for Moroccan Darija covering sentiment analysis, summarization, translation and transliteration, built for the Atlas-Chat model release.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategorymulti-task Moroccan Darija NLP suite: sentiment analysis, summarization, translation and transliteration
Page statusactive
Metricper-task-group metric (accuracy for sentiment; ROUGE/BERTScore/chrF for summarization; BLEU/chrF/TER/BERTScore for translation; BERTScore for transliteration)
Directionhigher_is_better
Unitmixed (%, points)
Dataset size31476
Dataset licenceODC-BY (top-level mirror licence); the dataset card explicitly warns this is a mixture with component-level licences that differ, some non-commercial -- see Dataset and licence.
PublisherMBZUAI-Paris, in a multi-institution collaboration with EMINES-UM6P, LINAGORA, KTH, AtlasIA and École Polytechnique

What it measures

DarijaBench bundles held-out test portions of pre-existing Darija datasets into four task groups: sentiment analysis (5 source datasets), summarization (1), translation (4, across six Darija-MSA/ English/French directions) and transliteration between Darija written in Arabic script and Arabizi written in Latin script (1). It does not test one skill; it is a fixed collection point for Darija-specific NLP evaluation, assembled because, as the authors state, standardized Darija benchmarks barely existed beforehand. Most component test portions are a 10% held-out split of their respective source dataset.

Task format

Format varies by group: sentiment is multiple-class classification (accuracy); summarization is free-text generation from a Darija article (ROUGE-1/2/L/Lsum, a BERTScore-style metric, and chrF); translation is free-text generation across six directions (BLEU, chrF, TER and a BERTScore-style metric); transliteration is free-text script conversion (the same BERTScore-style metric). All are zero-shot, single-turn generation tasks with no worked examples in the released harness configs.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub