3,482 native Modern Standard Arabic three-choice commonsense items across 13 countries; lm-eval group arab_culture, not translated English CSQA.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | native MSA cultural commonsense, 13 countries, 3-way MCQ |
| Page status | active |
| Metric | accuracy (acc; lm-eval also reports acc_norm); size-weighted mean across country tasks |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 3482 |
| Dataset licence | CC-BY-NC-SA-4.0 |
| Publisher | MBZUAI, with SDAIA, Al-Balqa Applied University, and Khalifa University |
ArabCulture tests whether a model can finish a culturally grounded commonsense statement written in Modern Standard Arabic. Native speakers from 13 countries wrote the items. They cover 12 daily-life domains and 54 subtopics (social norms, food, family, and similar), not school-exam facts. Each item is three-way multiple choice. The countries sit in four regions: Gulf (KSA, UAE, Yemen), Levant (Lebanon, Syria, Palestine, Jordan), North Africa (Tunisia, Algeria, Morocco, Libya), and Nile Valley (Egypt, Sudan). It is not a translation of English commonsense sets and not [arabic_mmlu](arabic_mmlu.md).
Three-option MCQ in MSA. lm-eval's arab_culture path ranks choice letters by log-likelihood (A/B/C or أ/ب/ج). A completion variant concatenates each option to the stem. Optional English vs Arabic prompt wrappers and optional region or country context are toggled with COUNTRY, REGION, and ARABIC environment variables.
No model card in ModelSpec reports this benchmark yet.