Anthropic publishes its Responsible Scaling Policy
AI Safety Levels borrowed the biosafety-lab naming scheme, and its rules would eventually pause deployment of any model reaching a level the company had not yet built safeguards for.
- Safety & alignment
- Major
Anthropic published its Responsible Scaling Policy, a self-imposed framework of technical and organisational commitments intended to keep the company’s own deployment decisions from outrunning its ability to manage the risk of catastrophic misuse. The policy borrowed its naming convention from biosafety laboratories, defining AI Safety Levels (ASL) that scale with a model’s demonstrated dangerous capabilities rather than with its size or benchmark scores.
ASL-1 covered systems posing no meaningful catastrophic risk, such as older language models. ASL-2 — where Anthropic placed Claude at the time — covered systems showing early signs of dangerous capability without the practical means to cause serious harm. ASL-3, not yet reached, would require materially stronger security around model weights and a commitment to halt deployment if red-teaming showed the model could meaningfully assist with catastrophic misuse, such as biological or chemical weapons development. ASL-4 and above were left undefined, with a commitment to specify them before any model approached that threshold. The policy also allowed for training itself to be paused if a model’s capabilities began outstripping the safeguards built to contain it, and required board approval, following consultation with Anthropic’s Long Term Benefit Trust, for any change to the framework.
Anthropic explicitly framed the approach as analogous to pre-market safety testing regimes in aviation and pharmaceuticals — testing before deployment rather than remediation afterward. The policy was voluntary and self-administered: nothing in it created an external body empowered to block a release if Anthropic judged its own safeguards adequate.
The RSP became the first of what turned into a genre. OpenAI published its own Preparedness Framework three months later, and Google DeepMind followed with its Frontier Safety Framework in May 2024 — each lab committing publicly to capability-triggered safety measures, and each leaving enforcement to itself.