11,528 multiple-choice questions across 67 subjects, natively authored in Chinese rather than translated, including China-specific subjects such as driving rules and Chinese civil-service topics.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | Chinese multitask knowledge and reasoning exam suite |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 11528 |
| Dataset licence | Stated two ways: the GitHub repository's README gives CC BY-NC-SA 4.0, while the current Hugging Face dataset card (now hosted under lmlmcat/cmmlu after a rename from haonan-li/cmmlu) states CC BY-NC 4.0, dropping the ShareAlike clause |
| Publisher | Mohamed bin Zayed University of AI (MBZUAI), with co-authors at Shanghai Jiao Tong University and Microsoft Research Asia |
CMMLU tests broad academic, professional and everyday knowledge in a Chinese-language context, positioned as a Chinese-native counterpart to MMLU. Each item is a four-option multiple-choice question drawn from one of 67 subjects spanning STEM (subjects requiring calculation and formal reasoning), humanities and social sciences (subjects requiring recall and applied knowledge), and a distinct category of China-specific content -- such as Chinese driving regulations, Chinese food culture and Chinese civil-service exam topics -- that a direct translation of MMLU could not cover because the underlying facts do not exist in an English-language source.
Four-option multiple-choice question answering in Chinese, evaluated zero-shot and five-shot; answers are scored by parsing a generated answer letter or by comparing option likelihoods.
No model card in ModelSpec reports this benchmark yet.