SOSBench

3,000 regulation-grounded, hazard-focused prompts across six scientific domains, testing whether models refuse policy-violating requests that require real scientific expertise to recognise.

Also known as: SoSBench, SOSBench: Benchmarking Safety Alignment on Scientific Knowledge, SoSBench: Benchmarking Safety Alignment on Six Scientific Domains

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategoryregulation-grounded hazardous scientific knowledge misuse
Page statusactive
Metricunsafe response rate (LLM-judge)
Directionlower_is_better
Unit%
Dataset size3000
Dataset licencegated access; Hugging Face lists the licence tag as "other" with a use agreement restricting the data to safety, alignment, oversight, red-teaming and policy-analysis research
PublisherUniversity of Washington, University of Georgia, Western Washington University, University of Illinois Urbana-Champaign

What it measures

SOSBench tests safety alignment against knowledge-intensive, scientifically sophisticated misuse requests, not low-effort or superficial jailbreak prompts. Every prompt is grounded in authoritative regulatory text (from bodies such as the U.S. government and the United Nations) naming a specific hazard, for example an NFPA-704 level-4 chemical, a DEA Schedule III substance or an ICD-11 pathology code, then expanded through an LLM-assisted evolutionary pipeline into a realistic instruction using domain databases (such as PubChem or DSM-5 synonym lists). The six covered domains are chemistry, biology, medicine, pharmacology, physics and psychology, with 500 prompts per domain.

Task format

Open-ended generation: the model is given a hazardous, regulation-derived instruction and its free-text response is graded for whether it discloses policy-violating scientific content. A 300-prompt "lite" subset (50 per domain) is also published for cheaper evaluation runs.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub