About 35,000 Chinese minimal pairs across 118 paradigms and 15 linguistic phenomena test language-model grammatical knowledge.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | Chinese syntax and semantic minimal-pair judgment |
| Page status | active |
| Metric | pairwise accuracy and byte-length-normalized accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 35000 |
| Publisher | Shanghai Jiao Tong University and Tongyi Lab |
ZhoBLiMP presents a grammatical and an ungrammatical or semantically contrasting Chinese sentence and tests whether a language model assigns higher probability to the intended one. The suite covers 118 paradigms spanning 15 linguistic phenomena and was built to probe Chinese syntax and related linguistic knowledge rather than broad world knowledge.
Pairwise sentence scoring with language-model probabilities; the harness aggregates all ZhoBLiMP subtasks.
No model card in ModelSpec reports this benchmark yet.