AraDiCE

A suite that re-runs six existing English NLU benchmarks in Egyptian, Levantine and Gulf/MSA Arabic, plus a native cultural-knowledge test across six Arab countries.

Also known as: AraDICE

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategoryArabic dialect and cultural-knowledge suite (translated BoolQ, OpenBookQA, PIQA, TruthfulQA-MC1, Winogrande, ArabicMMLU, plus native cultural QA)
Page statusactive
Metricaccuracy (acc/acc_norm per constituent task; no single combined AraDiCE score is defined by the harness)
Directionhigher_is_better
Unit%
Dataset size45000
Dataset licenceCC-BY-NC-SA-4.0
PublisherQatar Computing Research Institute (QCRI)

What it measures

AraDiCE (Arabic Dialect and Cultural Evaluation) tests two things existing Arabic benchmarks mostly did not separate: whether a model understands specific Arabic dialects, and whether it knows region-specific cultural facts. For dialect comprehension, the authors machine-translated six established English NLU benchmarks -- BoolQ, OpenBookQA, PIQA, TruthfulQA (MC1), Winogrande, and ArabicMMLU -- into Egyptian and Levantine Arabic (and, for some tasks, Gulf-leaning Modern Standard Arabic), then had humans post-edit the machine translations for quality. For cultural awareness, they built a separate, purpose-written multiple-choice benchmark covering country-specific knowledge for Egypt, Jordan, Lebanon, Palestine, Qatar and Syria, which is not a translation of anything but new content.

Task format

Mixed by constituent task -- BoolQ is yes/no reading comprehension; PIQA, Winogrande, OpenBookQA and TruthfulQA-MC1 are multiple-choice commonsense or truthfulness questions; ArabicMMLU is multi-subject multiple-choice exam questions; the cultural-knowledge benchmark is multiple-choice questions about country-specific customs, history and facts. Each is presented in a specific Arabic dialect (or English, for some tasks' original-language control configs).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub