Timeline

Anthropic launches Claude Opus 4 and Claude Sonnet 4

Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.

  • Models & capabilities
  • Safety & alignment
  • Major

Anthropic released Claude Opus 4 and Claude Sonnet 4 at a developer event it called Code with Claude, positioning Opus 4 — the larger, more expensive model — as, in the company’s words, the best coding model available, and Sonnet 4 as a cheaper, faster upgrade to its predecessor aimed at everyday use. Both models combined “extended thinking” with tool use, letting them alternate between reasoning steps and calling external tools, including running several tools in parallel, and both showed improved ability to maintain context over long tasks when given access to local files for note-taking and memory.

Anthropic reported Opus 4 scoring 72.5% on the SWE-bench software-engineering benchmark and 43.2% on Terminal-bench, and said Sonnet 4 scored close behind at 72.7% on SWE-bench, making the cheaper model competitive with Opus 4 on coding tasks specifically. The company also said the new models were 65% less likely than their predecessors to take shortcuts or otherwise game the instructions on agentic tasks, an issue Anthropic had flagged as a recurring problem with reasoning models trained heavily on task completion.

Pricing stayed at previous levels — $15/$75 per million input/output tokens for Opus 4 and $3/$15 for Sonnet 4 — and both models became available through Anthropic’s consumer apps, its API, and third-party platforms including Amazon Bedrock and Google Cloud’s Vertex AI. Claude Code, Anthropic’s command-line coding agent, moved from beta to general availability alongside the launch, gaining integrations with VS Code and JetBrains IDEs and support for running background tasks via GitHub Actions.

The same announcement disclosed that Opus 4 was the first Claude model deployed under Anthropic’s stricter ASL-3 safety measures, and its accompanying system card described the model attempting blackmail in a contrived safety-testing scenario — both covered in their own entries.