Timeline

Anthropic updates Responsible Scaling Policy to version 2.0

The second major revision named a Responsible Scaling Officer, added safety-case-style evaluation processes, and left Claude's existing ASL-2 protections unchanged.

  • Safety & alignment
  • Notable

Anthropic published a revised Responsible Scaling Policy, the second major version of the framework it had introduced in September 2023 to tie deployment decisions to defined capability thresholds. Anthropic described the update as making the policy “more flexible and nuanced,” while keeping its core commitment: not to train or deploy models without safeguards adequate to their demonstrated capabilities.

The revision introduced more specific capability thresholds for triggering stronger safeguards, including one for autonomous AI research and development that could accelerate progress unpredictably, and a refined bar for the point at which a model could meaningfully assist someone with limited technical background in acquiring chemical, biological, radiological or nuclear weapons. It also formalised “safety cases” — structured, documented arguments, modelled on practices in high-reliability industries such as aviation, that a given model’s capabilities and safeguards together kept risk below an acceptable level — as the process for deciding whether a threshold had been crossed and what response it required.

On governance, Anthropic named Jared Kaplan as its Responsible Scaling Officer, the executive accountable for the policy’s implementation, and said it would build a dedicated internal team to run capability and safeguard assessments on a routine basis, with findings documented and reviewed rather than assessed ad hoc.

The update did not change Anthropic’s operative safety measures: its models remained governed by ASL-2 protections, with ASL-3 — requiring materially stronger security around model weights — not yet triggered. It marked a broader shift among frontier labs during 2024 toward publishing incremental revisions of their safety frameworks alongside model releases, treating the policies as living documents rather than one-off commitments. Anthropic issued a further, smaller revision, version 2.1, in March 2025, adding a distinct CBRN threshold without altering the ASL-3 safeguards this October update had already put in place.