ArabCulture

3,482 native Modern Standard Arabic three-choice commonsense items across 13 countries; lm-eval group arab_culture, not translated English CSQA.

Also known as: Arab Culture, arab_culture, ArabicCulture, MBZUAI/ArabCulture

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorynative MSA cultural commonsense, 13 countries, 3-way MCQ
Page statusactive
Metricaccuracy (acc; lm-eval also reports acc_norm); size-weighted mean across country tasks
Directionhigher_is_better
Unit%
Dataset size3482
Dataset licenceCC-BY-NC-SA-4.0
PublisherMBZUAI, with SDAIA, Al-Balqa Applied University, and Khalifa University

What it measures

ArabCulture tests whether a model can finish a culturally grounded commonsense statement written in Modern Standard Arabic. Native speakers from 13 countries wrote the items. They cover 12 daily-life domains and 54 subtopics (social norms, food, family, and similar), not school-exam facts. Each item is three-way multiple choice. The countries sit in four regions: Gulf (KSA, UAE, Yemen), Levant (Lebanon, Syria, Palestine, Jordan), North Africa (Tunisia, Algeria, Morocco, Libya), and Nile Valley (Egypt, Sudan). It is not a translation of English commonsense sets and not [arabic_mmlu](arabic_mmlu.md).

Task format

Three-option MCQ in MSA. lm-eval's arab_culture path ranks choice letters by log-likelihood (A/B/C or أ/ب/ج). A completion variant concatenates each option to the stem. Optional English vs Arabic prompt wrappers and optional region or country context are toggled with COUNTRY, REGION, and ARABIC environment variables.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub