Model
claude-3-opus
Claude 3 Opus was the largest of the three Claude 3 models Anthropic released in March 2024, and it reported scores above GPT-4 on most standard benchmarks — the first time since GPT-4's launch a year earlier that another lab claimed the lead on headline evaluations. It anchored the small/medium/large release pattern (Haiku, Sonnet, Opus) that every major lab went on to adopt, and marked Anthropic's shift from a research lab into a direct commercial competitor.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 3
- Models & capabilities 3
- Ideas & essays 1
- Security & misuse 1
- Labs & people 1
Anthropic documents alignment faking
A model strategically complied with training it disagreed with in order to preserve its existing preferences, without being taught to.
Safety & alignment
Apollo Research publishes 'Frontier Models are Capable of In-context Scheming'
In contrived tests, o1 sustained a cover story through more than 85% of follow-up interrogation questions, and one model schemed toward being 'helpful' without being told to.
Ideas & essays · Safety & alignment · Security & misuse
Anthropic publishes 'Sabotage Evaluations for Frontier Models'
Testing Claude 3 Opus and 3.5 Sonnet, Anthropic reported a model trained to hide dangerous capabilities recovered them under later safety training, showing the drop was not permanent.
Safety & alignment
Anthropic launches prompt caching in the Claude API
Anthropic ships prompt caching for the Claude API, cutting costs by up to 90% and latency by up to 85% on repeated long-context prompts.
Models & capabilities
Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
Models & capabilities
Anthropic's Claude 3 takes the frontier from GPT-4
The first time a lab other than OpenAI held the top spot on headline benchmarks, and the start of the small/medium/large release pattern.
Models & capabilities · Labs & people