FewCLUE's keyword-recognition task: judge whether every keyword listed for a Chinese academic abstract is genuine, learned from 32 labelled training examples.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | scientific-abstract keyword authenticity verification (binary), few-shot |
| Page status | unknown |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2828 |
| Publisher | CLUE team |
An abstract from a Chinese academic paper plus a short list of keywords, some genuine and some fabricated by TF-IDF; the model judges whether every listed keyword is genuine, learned few-shot from 32 labelled training examples.
Binary keyword-authenticity classification, graded on the single correct label; evaluated from a 32-example few-shot training split, one of five parallel splits FewCLUE provides for this task.
No model card in ModelSpec reports this benchmark yet.