Timeline

Anthropic updates Responsible Scaling Policy to version 3.0

The policy now separates Anthropic's own commitments from industry-wide recommendations, adds a graded Frontier Safety Roadmap, and requires risk reports every three to six months.

  • Safety & alignment
  • Notable

Anthropic published version 3.0 of its Responsible Scaling Policy, its first major structural rewrite since the framework’s 2023 original. The policy governs how Anthropic decides which safety and security measures apply as its models cross defined capability thresholds (the ASL — AI Safety Level — scale).

The rewrite made three structural changes. It separated hard commitments Anthropic makes unilaterally from broader mitigations it recommends industry-wide, rather than presenting both as a single set of obligations. It replaced some fixed requirements with a “Frontier Safety Roadmap,” publicly declared, graded goals across security, alignment, safeguards and policy, allowing Anthropic to describe progress short of full compliance rather than being bound to all-or-nothing thresholds. And it introduced systematic risk reporting every three to six months, with third-party expert review triggered under specified circumstances.

The most consequential admission concerned safeguards for the highest capability tiers. Anthropic wrote that some ASL-4 and ASL-5 protections “might prove outright impossible to implement without collective action,” citing a RAND assessment that certain cybersecurity standards were “currently not possible” for any single company and would “likely require assistance from the national security community.” The document argued against defining those higher-tier safeguards in ways “easy to achieve,” saying that would undermine the policy’s purpose even where full compliance could not be guaranteed.

“We could have tried to address this by defining ASL-4 and ASL-5 safeguards in ways that made compliance easy to achieve — but this would undermine the intended spirit of the RSP.” — Anthropic, Responsible Scaling Policy version 3.0

The policy also confirmed Anthropic had implemented uniform ASL-3 deployment safeguards, including refined methods for blocking content of concern, since May 2025. Anthropic revised the framework again five weeks later, when RSP v3.1 took effect on 2 April 2026.