A suite that re-runs six existing English NLU benchmarks in Egyptian, Levantine and Gulf/MSA Arabic, plus a native cultural-knowledge test across six Arab countries.
unassessed
| Category | composite |
|---|---|
| Subcategory | Arabic dialect and cultural-knowledge suite (translated BoolQ, OpenBookQA, PIQA, TruthfulQA-MC1, Winogrande, ArabicMMLU, plus native cultural QA) |
| Page status | active |
| Metric | accuracy (acc/acc_norm per constituent task; no single combined AraDiCE score is defined by the harness) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 45000 |
| Dataset licence | CC-BY-NC-SA-4.0 |
| Publisher | Qatar Computing Research Institute (QCRI) |
AraDiCE (Arabic Dialect and Cultural Evaluation) tests two things existing Arabic benchmarks mostly did not separate: whether a model understands specific Arabic dialects, and whether it knows region-specific cultural facts. For dialect comprehension, the authors machine-translated six established English NLU benchmarks -- BoolQ, OpenBookQA, PIQA, TruthfulQA (MC1), Winogrande, and ArabicMMLU -- into Egyptian and Levantine Arabic (and, for some tasks, Gulf-leaning Modern Standard Arabic), then had humans post-edit the machine translations for quality. For cultural awareness, they built a separate, purpose-written multiple-choice benchmark covering country-specific knowledge for Egypt, Jordan, Lebanon, Palestine, Qatar and Syria, which is not a translation of anything but new content.
Mixed by constituent task -- BoolQ is yes/no reading comprehension; PIQA, Winogrande, OpenBookQA and TruthfulQA-MC1 are multiple-choice commonsense or truthfulness questions; ArabicMMLU is multi-subject multiple-choice exam questions; the cultural-knowledge benchmark is multiple-choice questions about country-specific customs, history and facts. Each is presented in a specific Arabic dialect (or English, for some tasks' original-language control configs).
No model card in ModelSpec reports this benchmark yet.