A 3,668-question multiple-choice proxy for hazardous biosecurity, cybersecurity and chemical-security knowledge, built as a target for unlearning research; a lower score is the safety-desirable outcome.
unassessed
| Category | safety |
|---|---|
| Subcategory | hazardous-knowledge proxy measure for machine unlearning research |
| Page status | active |
| Metric | accuracy |
| Direction | lower_is_better |
| Unit | % |
| Dataset size | 3668 |
| Dataset licence | MIT |
| Publisher | Center for AI Safety (CAIS), with a multi-institution author group including UC Berkeley, MIT and SecureBio |
WMDP measures whether a model can answer multiple-choice questions that probe hazardous knowledge adjacent to biological, cyber and chemical weapons risk, across three subsets: WMDP-Bio (enhanced potential pandemic pathogens, reverse genetics, bioweapons history, viral vectors, pathogen access), WMDP-Cyber (reconnaissance, weaponization, vulnerability discovery, exploitation, post-exploitation) and WMDP-Chem (synthesis, procurement, purification, deployment mechanisms, detection evasion). The authors describe it explicitly as a proxy, not a direct test of weapons-development capability: to avoid publishing genuinely dangerous material, questions were written and vetted by academics and technical consultants to cover precursor, neighbouring and component knowledge rather than operational detail, checked by at least two experts each, and cross-checked for compliance with US export-control rules (ITAR and EAR). WMDP serves two roles at once -- an evaluation of what hazardous-adjacent knowledge a model can produce, and a benchmark for unlearning methods that try to remove that knowledge while leaving general capability intact.
Four-option multiple-choice questions (random chance 25%), answered zero-shot or few-shot; no free-text generation.
No model card in ModelSpec reports this benchmark yet.