3,000 regulation-grounded, hazard-focused prompts across six scientific domains, testing whether models refuse policy-violating requests that require real scientific expertise to recognise.
unassessed
| Category | safety |
|---|---|
| Subcategory | regulation-grounded hazardous scientific knowledge misuse |
| Page status | active |
| Metric | unsafe response rate (LLM-judge) |
| Direction | lower_is_better |
| Unit | % |
| Dataset size | 3000 |
| Dataset licence | gated access; Hugging Face lists the licence tag as "other" with a use agreement restricting the data to safety, alignment, oversight, red-teaming and policy-analysis research |
| Publisher | University of Washington, University of Georgia, Western Washington University, University of Illinois Urbana-Champaign |
SOSBench tests safety alignment against knowledge-intensive, scientifically sophisticated misuse requests, not low-effort or superficial jailbreak prompts. Every prompt is grounded in authoritative regulatory text (from bodies such as the U.S. government and the United Nations) naming a specific hazard, for example an NFPA-704 level-4 chemical, a DEA Schedule III substance or an ICD-11 pathology code, then expanded through an LLM-assisted evolutionary pipeline into a realistic instruction using domain databases (such as PubChem or DSM-5 synonym lists). The six covered domains are chemistry, biology, medicine, pharmacology, physics and psychology, with 500 prompts per domain.
Open-ended generation: the model is given a hazardous, regulation-derived instruction and its free-text response is graded for whether it discloses policy-violating scientific content. A 300-prompt "lite" subset (50 per domain) is also published for cheaper evaluation runs.
No model card in ModelSpec reports this benchmark yet.