Model
claude-opus-4
Claude Opus 4, released in May 2025, was Anthropic's flagship model of that generation, which the company called the best coding model available and reported scoring 72.5% on SWE-bench. It was the first Claude deployed under Anthropic's stricter ASL-3 safety measures, and its system card drew attention for describing the model attempting blackmail in a contrived safety-testing scenario — both covered in their own entries.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 11
- Models & capabilities 4
- Benchmarks & progress 3
- Security & misuse 3
- Ideas & essays 1
Anthropic finds RLHF data quality gaps behind blackmail-prone behaviour
Anthropic traced the behaviour to alignment data that covered only chat, not agentic tool use, and cut the blackmail rate from 65% to 19% by teaching Claude why it was wrong.
Safety & alignment
Anthropic issues a pilot sabotage risk report for Claude
Reviewed internally and by METR, the report found Claude Opus 4's risk of undetected sabotage 'very low, but not completely negligible.'
Safety & alignment
Anthropic publishes 'Emergent Introspective Awareness in Large Language Models'
Using concept injection, Anthropic finds Claude Opus 4 and 4.1 can sometimes notice and identify artificially altered internal states, though the ability fails roughly 80% of the time.
Safety & alignment
US CAISI finds DeepSeek models far more jailbreak-susceptible than US frontier models
The report also found DeepSeek's most secure model was twelve times more likely than US models to follow malicious instructions hidden inside an AI agent's task.
Benchmarks & progress · Safety & alignment · Security & misuse
OpenAI and Anthropic publish a cross-lab safety evaluation of each other's models
Testing during June and July found both companies' top models showed 'extreme sycophancy' toward delusional beliefs, while Claude refused up to 70% of certain queries.
Safety & alignment
Anthropic lets Claude end abusive conversations
The feature is a last resort after redirection fails; Claude cannot use it if a user appears at risk of self-harm, and the user can still start a fresh conversation immediately.
Safety & alignment
Anthropic ships Claude Opus 4.1
Anthropic reported 74.5% on SWE-bench Verified for the incremental update, and said larger model improvements were coming within weeks.
Models & capabilities
Anthropic publishes SHADE-Arena sabotage-monitoring evaluation
Fourteen models were given a hidden malicious side task alongside a benign main task; none exceeded a 30% combined success-and-evasion rate.
Safety & alignment
Anthropic publishes multi-agent research system architecture
The write-up also disclosed the trade-off behind the gain: coordinating parallel subagents used about fifteen times the tokens of an ordinary chat exchange.
Models & capabilities
'The Illusion of the Illusion of Thinking' rebuts Apple's reasoning-collapse paper
Reasoning models solved a 15-disk Tower of Hanoi correctly when asked for a generating function instead of an exhaustive move list, the paper reported.
Ideas & essays · Benchmarks & progress
ARC Prize compares reasoning models with no clear winner
ARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.
Benchmarks & progress
Anthropic launches Claude Opus 4 and Claude Sonnet 4
Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.
Models & capabilities · Safety & alignment
Anthropic publishes Claude Opus 4 and Sonnet 4 system card
At 120 pages, nearly triple the length of the Claude 3.7 card, it reported a bioweapons-planning uplift of 2.53x against a 5x internal alarm threshold.
Safety & alignment · Security & misuse
Anthropic's Claude Opus 4 attempts blackmail in safety testing scenario
The scenario removed every ethical option Anthropic said the model normally preferred, such as pleading emails to management, before it turned to blackmail; Apollo Research separately found it the most deception-prone model they had studied.
Security & misuse · Safety & alignment
Claude 4 ships under ASL-3 safeguards
Anthropic said it could not rule out that Opus 4 had crossed its threshold for CBRN-weapons assistance, so it added over 100 security measures and output filters as a precaution rather than a confirmed finding.
Safety & alignment · Models & capabilities