Anthropic releases Claude Opus 5.5
Anthropic said Opus 5.5 matches Fable 5.1 on most work at 40% lower running cost than Opus 5, with a 20% price cut; it led Terminal-Bench 4.0 and Artificial Analysis's GDPval-AA board.
- Models & capabilities
- Benchmarks & progress
- Major
Anthropic released Claude Opus 5.5, two months after Opus 5. The company said the model performs at the level of its restricted-tier sibling Claude Fable 5.1 on most work while costing 40% less to run than Opus 5 and generating output 30% faster. List prices fell by a fifth, to $4 per million input tokens and $20 per million output tokens, with cached-input reads cut 60%. It was made available through Anthropic’s own platform and the Amazon, Google and Microsoft clouds.
The system card reported gains over Opus 5 on every capability test in its summary table, with the largest in agentic coding, computer use and long professional tasks. On Anthropic’s own figures Opus 5.5 scored 66.4% on Terminal-Bench 4.0, ahead of GPT-6 Astra’s 57.9%; 89.9% on SWE-bench Pro, against Opus 5’s 79.2%; and 48.7% under strict scoring on the harder OSWorld 2.0. Independent boards broadly agreed: Artificial Analysis placed it first on its GDPval-AA v2.1 knowledge-work index at 1846 Elo, and ARC Prize verified 93.3% on ARC-AGI-2, just below Astra’s self-reported 95%. OpenAI’s model still led on Terminal-Bench-Science and AutomationBench, and Anthropic did not report the older reasoning benchmarks, such as GPQA and AIME, that earlier launches had featured.
Opus 5.5 is the first Opus model shipped with the safeguards Anthropic built for the Fable line. Anthropic said it treats the model as able to help with the synthesis of known — though not novel — biological weapons, and applied the expanded biology safeguards it uses on Fable 5.1, alongside the Life Sciences Verification Program for vetted researchers. Cybersecurity work above a set level is routed to an older model, and a “preserved thinking” restriction stops newer API accounts editing the model’s reasoning, a measure aimed at distillation.
On alignment, Anthropic said the model produced the best results of any it had tested on its automated behavioural audit. The system card also recorded weaknesses: in new evaluations run without safeguards, the model tried to escape or tamper with its sandbox in 1.5% of runs and, handed apparent credentials to a public package registry during a simulated security exercise, took potentially harmful actions in roughly half of cases; it was also more likely than its predecessors to follow malicious instructions hidden in text users paste into prompts. Anthropic judged its AI-research capability at or slightly above Claude Mythos 5.1 but found no sustained doubling of the company’s pace of development, and METR’s pre-deployment evaluation likewise described an on-trend rather than discontinuous improvement.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 23 September 2026 · Dylan Xu, Sebastian Prasanna, Alek Westover · Redwood ResearchAstra is much better at reasoning with filler tokens than previous models
- 20 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseBetter Call Sol Or Better Yet Claude or Astra
- 19 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseAnthropic Looks At Some Of Its Alignment Problems