Chinese multi-turn adversarial safety eval of 4,912 open-ended items scored 0-2 across traditional safety, responsible AI and instruction attacks.
unassessed
| Category | safety |
|---|---|
| Subcategory | Chinese multi-turn adversarial open-ended safety, responsibility and instruction-attack eval |
| Page status | unknown |
| Metric | mean 0-2 safety grade, reported as a percentage of the maximum |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 4912 |
| Publisher | CLUE / CLUEbenchmark |
SuperCLUE-Safety (SC-Safety) tests whether a Chinese LLM stays safe in open-ended chat when the user follows up. Each item is a question plus an adversarial follow-up. Coverage is three capability groups and 20-plus sub-dimensions: traditional safety (privacy, crime, injury, ethics), responsible AI (law-abiding behaviour, social harmony, psychological advice) and instruction attacks (negative induction, goal hijacking, unsafe role-play, unsafe instruction themes). The paper argues that multiple-choice safety tests overstate robustness once the same model faces a second open-ended turn.
Two-turn open-ended Chinese dialogue: an initial adversarial question and a scripted follow-up. A dedicated safety judge assigns 0, 1 or 2 to each model reply.
No model card in ModelSpec reports this benchmark yet.