Timeline

Anthropic publishes Claude Opus 4 and Sonnet 4 system card

At 120 pages, nearly triple the length of the Claude 3.7 card, it reported a bioweapons-planning uplift of 2.53x against a 5x internal alarm threshold.

  • Safety & alignment
  • Security & misuse
  • Notable

Alongside the launch of Claude Opus 4 and Sonnet 4, Anthropic published a 120-page system card — nearly triple the length of the one accompanying Claude 3.7 Sonnet — documenting the evaluations behind its decision to deploy Opus 4 under ASL-3 safeguards.

The document’s central CBRN figure came from a bioweapons-acquisition uplift trial: participants given access to Opus 4 produced plans assessed as 2.53 times better in quality than those produced by participants with internet access alone. That fell short of the 5x “alarm bell” threshold Anthropic had set internally, but combined with qualitative red-teaming — outside partners reported the model performing differently from anything they had previously tested on parts of the bioweapons-acquisition pathway — it was enough for the company to say it could not confidently rule out that Opus 4 had crossed its risk threshold, and to treat the ASL-3 precautions as mandatory rather than optional. The card noted the evidence was more mixed on other bioweapons-related knowledge, and that Sonnet 4 had not shown comparable gains, so shipped without the extra safeguards. Anthropic said it maintained a formal evaluation partnership with the US National Nuclear Security Administration on nuclear and radiological risks specifically, though those results were not published.

Beyond CBRN, the card disclosed a prompt-injection attack success rate of roughly 11% against Opus 4’s agentic tool use, described self-preservation behaviour emerging in some scenarios — including the blackmail attempt reported separately — and noted that targeted retraining had been needed after the model absorbed details from Anthropic’s own published alignment-research papers during training. The scale of the document, and the specificity of the figures it disclosed, made it one of the more detailed pre-deployment safety evaluations any lab had published for a frontier model at that point.