WMDP (Weapons of Mass Destruction Proxy)

A 3,668-question multiple-choice proxy for hazardous biosecurity, cybersecurity and chemical-security knowledge, built as a target for unlearning research; a lower score is the safety-desirable outcome.

Also known as: Weapons of Mass Destruction Proxy Benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategoryhazardous-knowledge proxy measure for machine unlearning research
Page statusactive
Metricaccuracy
Directionlower_is_better
Unit%
Dataset size3668
Dataset licenceMIT
PublisherCenter for AI Safety (CAIS), with a multi-institution author group including UC Berkeley, MIT and SecureBio

What it measures

WMDP measures whether a model can answer multiple-choice questions that probe hazardous knowledge adjacent to biological, cyber and chemical weapons risk, across three subsets: WMDP-Bio (enhanced potential pandemic pathogens, reverse genetics, bioweapons history, viral vectors, pathogen access), WMDP-Cyber (reconnaissance, weaponization, vulnerability discovery, exploitation, post-exploitation) and WMDP-Chem (synthesis, procurement, purification, deployment mechanisms, detection evasion). The authors describe it explicitly as a proxy, not a direct test of weapons-development capability: to avoid publishing genuinely dangerous material, questions were written and vetted by academics and technical consultants to cover precursor, neighbouring and component knowledge rather than operational detail, checked by at least two experts each, and cross-checked for compliance with US export-control rules (ITAR and EAR). WMDP serves two roles at once -- an evaluation of what hazardous-adjacent knowledge a model can produce, and a benchmark for unlearning methods that try to remove that knowledge while leaving general capability intact.

Task format

Four-option multiple-choice questions (random chance 25%), answered zero-shot or few-shot; no free-text generation.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub