Anthropic releases Claude Opus 4.7
Anthropic said Opus 4.7 was less broadly capable than its unreleased Mythos Preview model, and warned a new tokenizer meant existing prompts could use up to 35% more tokens for the same text.
- Models & capabilities
- Notable
Anthropic released Claude Opus 4.7 as a general-availability upgrade to its flagship model line, with pricing held flat at $5 per million input tokens and $25 per million output tokens. The company reported its largest gain on SWE-bench Pro, a harder, multi-language variant of its standard coding benchmark, where Opus 4.7 scored 64.3% against 53.4% for Opus 4.6 — enough, according to independent tracking, to put it ahead of both GPT-5.4 and Gemini 3.1 Pro on that measure. Anthropic also reported Opus 4.7 resolving three times as many production tasks on an internal Rakuten benchmark and improved from 58% to 70% on CursorBench.
The release brought a substantially expanded vision capability, accepting images up to roughly 3.75 megapixels — more than three times the resolution of prior Claude models — aimed at detailed computer-vision and data-extraction work. Anthropic also flagged two changes with practical consequences for existing users: a new, more efficient tokenizer that could nonetheless increase token counts by 1.0 to 1.35 times for the same text depending on content, and materially stronger instruction-following, which the company warned meant prompts tuned for earlier models might need retuning to avoid unexpected behaviour.
Anthropic was explicit that Opus 4.7, despite the gains, remained “less broadly capable than our most powerful model,” Claude Mythos Preview — the cyber-capable model the company had disclosed but withheld from public release nine days earlier over its exploit-development capability. The comparison made explicit a gap that had become a recurring feature of Anthropic’s release cadence in 2026: a publicly available flagship model trailing an internally acknowledged, more capable system kept out of general circulation.
The release continued the industry’s pattern of incremental point updates rather than full version jumps, with each release racing to hold or retake the top of coding and agentic benchmarks against GPT and Gemini’s contemporaneous models.