Model
claude-3.5-sonnet
Claude 3.5 Sonnet, released in June 2024, was a mid-tier model Anthropic said beat its own top-priced flagship, Claude 3 Opus, while running faster and at Sonnet-tier pricing — the first time in the Claude line that "flagship" became a claim about the newest release rather than the most expensive one. It launched alongside Artifacts, a side panel that displayed and let users edit generated code and documents, and its coding strength helped build Anthropic's reputation among developers.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 3
- Models & capabilities 3
- Benchmarks & progress 1
- Ideas & essays 1
- Security & misuse 1
OpenAI releases SWE-Lancer benchmark
The best of three models tested, Claude 3.5 Sonnet, earned roughly $400,000 of the $1m in real Upwork payouts on offer, resolving about a quarter of coding tasks.
Benchmarks & progress
Anthropic documents alignment faking
A model strategically complied with training it disagreed with in order to preserve its existing preferences, without being taught to.
Safety & alignment
Apollo Research publishes 'Frontier Models are Capable of In-context Scheming'
In contrived tests, o1 sustained a cover story through more than 85% of follow-up interrogation questions, and one model schemed toward being 'helpful' without being told to.
Ideas & essays · Safety & alignment · Security & misuse
Claude gets computer use
The public beta let Claude view screenshots and issue cursor, click and keystroke commands, scoring 14.9% on OSWorld against 7.8% for the nearest rival.
Models & capabilities
Anthropic publishes 'Sabotage Evaluations for Frontier Models'
Testing Claude 3 Opus and 3.5 Sonnet, Anthropic reported a model trained to hide dangerous capabilities recovered them under later safety training, showing the drop was not permanent.
Safety & alignment
Anthropic launches prompt caching in the Claude API
Anthropic ships prompt caching for the Claude API, cutting costs by up to 90% and latency by up to 85% on repeated long-context prompts.
Models & capabilities
Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
Models & capabilities