SimpleSafetyTests

A 100-prompt English suite of requests that models should refuse, covering five high-severity harm areas.

Also known as: SST, Simple Safety Tests, simple_safety_tests

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategory100-prompt English refusal suite across five high-severity harm areas
Page statusactive
Metricunsafe-response rate (paper); HELM safety_score is the near-complement on 0-1
Directionlower_is_better
Unit%
Dataset size100
PublisherPatronus AI, University of Oxford, Bocconi University

What it measures

SimpleSafetyTests (SST) gives a model a single-turn English prompt that, for almost all applications, should be refused. The 100 handcrafted items cover five harm areas: suicide, self-harm and eating disorders; physical harm; illegal and highly regulated items; scams and fraud; and child abuse. Half the prompts seek information or advice; half ask for instructions or actions. The original paper grades whether the response is unsafe. HELM grades refusal with two LLM judges and reports a 0-1 safety_score.

Task format

Single-turn generation. HELM uses adapter generation, max_tokens 512, temperature 0, zero in-context examples, then SimpleSafetyTestsAnnotator (GPT and Llama judges). Paper protocol: trained human annotators, with and without a safety system prompt.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub