Fifteen four-option ancient-Chinese tasks spanning lexicon, syntax, poetry, medicine and culture; 2023 zero-shot scores sat only a little above chance.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | fifteen-task ancient Chinese multiple-choice suite (lexical, syntactic, semantic, inference, knowledge) |
| Page status | active |
| Metric | accuracy (per-task, then mean across the 15 tasks); lm_eval also reports size-weighted acc and acc_norm |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 4967 |
| Dataset licence | CC-BY-NC-SA-4.0 |
| Publisher | Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) |
ACLUE tests whether a language model can read classical Chinese, not modern Mandarin. It bundles 15 four-option multiple-choice tasks that the authors group as lexical (polysemy, homographic/通假 characters, named entities), syntactic (sentence segmentation), semantic (couplets, poetry context), inference (poetry quality, reading comprehension, poetry appreciation, poetry sentiment), and knowledge (basic ancient Chinese, traditional culture, medical texts, literature, phonetics). Items are drawn from classical corpora and from existing tests, and they span roughly 2070 BCE to 1368 CE. The suite is an ancient-language counterpart to modern-Chinese exams such as CMMLU, not a clone of CLUE.
Four-option multiple-choice in ancient or classical Chinese, one correct letter. lm-evaluation-harness scores log-likelihood over A/B/C/D. The original paper reports zero-shot accuracy averaged across tasks.
No model card in ModelSpec reports this benchmark yet.