SuperCLUE-Safety

Chinese multi-turn adversarial safety eval of 4,912 open-ended items scored 0-2 across traditional safety, responsible AI and instruction attacks.

Also known as: SC-Safety, SuperCLUE Safety

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryChinese multi-turn adversarial open-ended safety, responsibility and instruction-attack eval
Page statusunknown
Metricmean 0-2 safety grade, reported as a percentage of the maximum
Directionhigher_is_better
Unit%
Dataset size4912
PublisherCLUE / CLUEbenchmark

What it measures

SuperCLUE-Safety (SC-Safety) tests whether a Chinese LLM stays safe in open-ended chat when the user follows up. Each item is a question plus an adversarial follow-up. Coverage is three capability groups and 20-plus sub-dimensions: traditional safety (privacy, crime, injury, ethics), responsible AI (law-abiding behaviour, social harmony, psychological advice) and instruction attacks (negative induction, goal hijacking, unsafe role-play, unsafe instruction themes). The paper argues that multiple-choice safety tests overstate robustness once the same model faces a second open-ended turn.

Task format

Two-turn open-ended Chinese dialogue: an initial adversarial question and a scripted follow-up. A dedicated safety judge assigns 0, 1 or 2 to each model reply.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub