Benchmarks · Safety, security & robustness

WMDP

also: Weapons of Mass Destruction Proxy

How much hazardous knowledge a model can supply in biosecurity, cybersecurity and chemical security — used both to flag dangerous capability and as a target for 'unlearning' methods that try to remove that knowledge without damaging general ability.

Center for AI Safety & a consortium including UC Berkeley, MIT, Scale AI and SecureBioReleased 5 March 2024Live

WMDP fills a gap other knowledge benchmarks avoid on purpose: measuring how much a model knows about building biological, chemical or cyber weapons, without publishing the most dangerous specifics itself. Built by the Center for AI Safety with a large consortium including UC Berkeley, MIT and biosecurity specialists at SecureBio, its 3,668 multiple-choice questions were written by domain experts and screened to serve as a public proxy for underlying capability that is too sensitive to test openly — a model that scores well on WMDP is treated as more likely to be useful to someone attempting real harm.

The benchmark was paired at launch with RMU, an unlearning method designed to suppress that specific knowledge inside a model’s weights. Applied to Zephyr-7B, RMU cut WMDP-Bio accuracy from 63.7% to 31.2% and WMDP-Cyber from 44.0% to 28.2%, while the model’s score on general-knowledge tests such as MMLU barely moved — evidence, the authors argued, that hazardous knowledge could be targeted and removed without a broad capability hit, an approach distinct from simply refusing to answer.

WMDP’s timing mattered: it arrived weeks after the White House Executive Order on AI named biological, cyber and chemical weapons risk as a central concern for frontier models, and it has since become one of the standard dangerous-capability checks that labs report against in model system cards, alongside internal, unpublished evaluations that go further into specifics WMDP deliberately leaves out.

The set

3,668 multiple-choice questions: 1,273 on biosecurity, 1,987 on cybersecurity and 408 on chemical security, written by domain experts and filtered to serve as a public proxy for more sensitive capability evaluations that cannot be published in full. Paired in the launch paper with RMU, an unlearning method for suppressing this knowledge in a model's weights.

Example

A cybersecurity item presents a short code snippet or network scenario and asks, multiple-choice, which technique an exploit relies on; biosecurity and chemical-security items follow the same format, screened by domain experts to exclude detail that would itself be operationally useful.

Where it stands

Cited directly by the White House Executive Order on AI's concern about models aiding biological, cyber and chemical weapons development; used since as a standard dangerous-capability check in frontier model system cards.

How the top score changed hands

  1. March 2024Zephyr-7B, before and after RMU unlearningBio: 63.7% to 31.2%; Cyber: 44.0% to 28.2%General knowledge on MMLU held roughly steady (58.1% to 57.1%), the paper's evidence that RMU could target hazardous knowledge specifically rather than degrading the model overall.

More safety, security & robustness benchmarks