Anthropic ships Claude Sonnet 4.5
Anthropic reported 77.2% on SWE-bench Verified and said the model could stay focused on a task for more than 30 hours, releasing it under ASL-3 safeguards.
- Models & capabilities
- Notable
Anthropic released Claude Sonnet 4.5, calling it the strongest coding model it had shipped and the model best suited to building complex, long-running agents. The company reported a score of 77.2% on SWE-bench Verified and 61.4% on OSWorld, a computer-use benchmark, up sharply from the 42.2% Claude Sonnet 4 had scored on the same test four months earlier. Anthropic said the model could sustain focus on complex, multi-step tasks for more than 30 hours without losing coherence, aided by new memory and context-editing tools designed to manage long agentic sessions.
Pricing was unchanged from Sonnet 4, at $3 per million input tokens and $15 per million output tokens, positioning the release as a capability upgrade at the existing price point rather than a new tier. Anthropic also described gains on reasoning, mathematics and computer-use tasks, and reported strong domain-specific performance in areas including finance, law, medicine and STEM evaluations.
The release was accompanied by a system card describing the model as released under Anthropic’s ASL-3 safeguards, including classifiers aimed at CBRN-related misuse, and by claims — based on the company’s own interpretability and behavioural evaluations — that Sonnet 4.5 showed reduced rates of deception, sycophancy and power-seeking behaviour relative to earlier Claude models. Anthropic described the model as its “most aligned frontier model yet,” a characterisation that, as with similar claims from other labs, rested on the company’s own evaluation suite rather than an independently audited standard.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.