ACLUE (Ancient Chinese Language Understanding Evaluation)

Fifteen four-option ancient-Chinese tasks spanning lexicon, syntax, poetry, medicine and culture; 2023 zero-shot scores sat only a little above chance.

Also known as: Ancient Chinese Language Understanding Evaluation

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategoryfifteen-task ancient Chinese multiple-choice suite (lexical, syntactic, semantic, inference, knowledge)
Page statusactive
Metricaccuracy (per-task, then mean across the 15 tasks); lm_eval also reports size-weighted acc and acc_norm
Directionhigher_is_better
Unit%
Dataset size4967
Dataset licenceCC-BY-NC-SA-4.0
PublisherMohamed bin Zayed University of Artificial Intelligence (MBZUAI)

What it measures

ACLUE tests whether a language model can read classical Chinese, not modern Mandarin. It bundles 15 four-option multiple-choice tasks that the authors group as lexical (polysemy, homographic/通假 characters, named entities), syntactic (sentence segmentation), semantic (couplets, poetry context), inference (poetry quality, reading comprehension, poetry appreciation, poetry sentiment), and knowledge (basic ancient Chinese, traditional culture, medical texts, literature, phonetics). Items are drawn from classical corpora and from existing tests, and they span roughly 2070 BCE to 1368 CE. The suite is an ancient-language counterpart to modern-Chinese exams such as CMMLU, not a clone of CLUE.

Task format

Four-option multiple-choice in ancient or classical Chinese, one correct letter. lm-evaluation-harness scores log-likelihood over A/B/C/D. The original paper reports zero-shot accuracy averaged across tasks.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub