Anthropic and US National Nuclear Security Administration build a nuclear-content classifier
The classifier, co-developed with the Department of Energy's NNSA and already running on live Claude traffic, reached 96% accuracy in preliminary testing.
- Safety & alignment
- Security & misuse
- Minor
Anthropic announced that it had worked with the US Department of Energy’s National Nuclear Security Administration to build a classifier able to distinguish conversations about nuclear weapons that raised proliferation concerns from benign discussion of nuclear topics such as energy policy or physics. The NNSA, drawing on its classified expertise in weapons-relevant material, helped Anthropic define what dangerous content looked like without either party needing to expose classified information to the other’s systems directly.
Anthropic said the resulting classifier reached 96% accuracy in preliminary testing and that early results from running it against real Claude conversations suggested it performed well outside the lab. The company said it had already deployed the classifier on live Claude traffic as part of its broader misuse-detection system, rather than treating the collaboration as a research exercise alone.
Anthropic framed the work as a template, saying it intended to share the approach with the Frontier Model Forum so other AI developers could build comparable safeguards. The partnership was one of a small number of documented cases in which a US national-security agency and a frontier AI lab worked jointly on a specific technical safeguard against catastrophic misuse, rather than the more common pattern of policy engagement or evaluation access, and it followed a broader run of 2025 announcements in which frontier labs sought to demonstrate cooperation with government on nuclear, biological and chemical risks specifically.