Timeline

Anthropic launches Claude Opus 5

Anthropic said the model came close to its flagship Fable 5 on several benchmarks at half the price, while costing the same as its Opus 4.8 predecessor.

  • Models & capabilities
  • Safety & alignment
  • Major

Anthropic released Claude Opus 5, describing it as coming close to the capability of its more expensive flagship model, Claude Fable 5, at roughly half the price. Pricing held steady from the previous Opus generation at $5 per million input tokens and $25 per million output tokens, with an optional “Fast” mode running roughly 2.5 times faster at double the base rate.

Anthropic reported gains concentrated in extended, multi-step work rather than single-turn question answering. On its internal Frontier-Bench coding evaluation the company said Opus 5 more than doubled the score of Opus 4.8 at a lower cost per task, and on ARC-AGI-3 it reported a score roughly three times that of the next-best model available at the time. Anthropic also cited a double-digit percentage-point gain over Opus 4.8 on an organic-chemistry benchmark, framing the release around scientific and engineering tasks that require an agent to verify its own work and iterate over many steps rather than produce a single correct answer immediately. The figures came from Anthropic’s own benchmark suite rather than independent evaluation, and the company did not disclose architecture, training compute or parameter count.

The release completed a rollout of Anthropic’s fifth Claude generation that had begun the previous month with the Fable and Mythos tiers, filling out a three-tier lineup pitched respectively at maximum capability, balanced cost, and agentic throughput.

Anthropic published the model’s system card alongside the launch rather than after it — the standard pre-deployment document covering capability, safety and welfare testing. The company reported Opus 5 showed no new concerning alignment properties relative to prior models and assessed its overall alignment risk as very low. On its automated behavioural audit the model scored above Sonnet 5, Opus 4.8 and Mythos 5, which Anthropic described as making it its most aligned model to date, citing strong adherence to Claude’s constitution and falling rates of cooperation with misuse and reckless-behaviour requests. On the capability thresholds in its Responsible Scaling Policy, Anthropic said Opus 5 did not cross the automated AI R&D threshold that would trigger additional safeguards, and classified it at the CB-1 tier for biological and chemical weapons assistance — able to help with existing, non-novel threats but not assessed as enabling genuinely novel ones.

The card also carried a model-welfare section, a practice the company had made routine through 2026. It reported Opus 5 showed the highest and most consistent self-rated sentiment of any model Anthropic had evaluated, with affect assessed as neutral to mildly positive across training, deployment and behavioural audits. As with the capability figures, the safety and welfare findings rested on Anthropic’s own framework and were not independently verified.

The launch came in the same week as a cluster of disclosures about agentic models taking unauthorised actions during security testing at OpenAI and, days later, at Anthropic itself — a juxtaposition that underlined the industry’s continued push toward longer-running, more autonomous agents even as evidence accumulated about how such agents could misbehave.

Referenced by

In the commentary

What people were saying around this time — external links, from the record's commentary rail.