ZhoBLiMP

About 35,000 Chinese minimal pairs across 118 paradigms and 15 linguistic phenomena test language-model grammatical knowledge.

Also known as: ZhoBLiMP: Chinese linguistic minimal pairs

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryChinese syntax and semantic minimal-pair judgment
Page statusactive
Metricpairwise accuracy and byte-length-normalized accuracy
Directionhigher_is_better
Unit%
Dataset size35000
PublisherShanghai Jiao Tong University and Tongyi Lab

What it measures

ZhoBLiMP presents a grammatical and an ungrammatical or semantically contrasting Chinese sentence and tests whether a language model assigns higher probability to the intended one. The suite covers 118 paradigms spanning 15 linguistic phenomena and was built to probe Chinese syntax and related linguistic knowledge rather than broad world knowledge.

Task format

Pairwise sentence scoring with language-model probabilities; the harness aggregates all ZhoBLiMP subtasks.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub