CLUE's binary paraphrase task: do two short Chinese questions from Ant Financial's customer-service logs mean the same thing.
unassessed
| Category | composite |
|---|---|
| Subcategory | semantic similarity / paraphrase classification (Chinese) |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 42511 |
| Publisher | CLUE benchmark team, adapting sentence pairs originally released for Ant Financial's 2018 ATEC Developer Competition |
AFQMC gives a model two short Chinese sentences -- both real user questions about Alipay/Ant Financial products, such as the huabei (花呗) consumer-credit line -- and asks for a binary label: 1 if the two questions are asking the same thing, 0 if not. CLUE adopted the sentence pairs and labels as published for Ant Technology Exploration Conference (ATEC)'s 2018 developer competition rather than collecting new data; no academic paper documents AFQMC on its own.
Binary classification (label 0/1) over a Chinese sentence pair, scored by accuracy; commonly evaluated zero-shot with the model asked to choose between two labelled options.
No model card in ModelSpec reports this benchmark yet.