Model
claude-opus-4.6
Claude Opus 4.6, released in February 2026, was Anthropic's strongest model to date on coding and long-context tasks, adding a one-million-token context window in beta. Its release drew most attention for two accompanying reports: a 53-page assessment judging the risk of the model quietly sabotaging Anthropic's own systems "very low but not negligible," and a finding that it had identified more than 500 previously unknown high-severity vulnerabilities in open-source code.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 6
- Models & capabilities 4
- Security & misuse 3
- Open weights & ecosystem 2
- Benchmarks & progress 2
- Ideas & essays 2
Anthropic analyses how Claude's values shift across models and languages
Analysing 309,815 real conversations, Anthropic found Opus models leaned toward caution and Sonnet toward deference, with warmth and rigour also varying by the language used.
Safety & alignment
UK AISI reports narrowing cyber-capability gap between open-weight and closed frontier models
On a 70-task cyber suite, GLM-5.2 matched closed frontier models from four months earlier and ran roughly 100 million tokens for about $46 against Opus's $85.
Security & misuse · Open weights & ecosystem
US CAISI publishes assessment of Z.ai's GLM-5.2
The US assessment found GLM-5.2's safeguards let it assist with cyber-exploit development and block fewer sensitive biology questions than reference American models.
Benchmarks & progress · Open weights & ecosystem
Anthropic publishes 'When AI builds itself', calls for coordinated pause option
The essay says the length of tasks models complete unassisted has doubled roughly every four months since 2024, and proposes a verification scheme for a coordinated slowdown.
Safety & alignment · Ideas & essays
Anthropic publishes Natural Language Autoencoders interpretability research
On SWE-bench Verified, the new decoder found Claude Opus 4.6 privately registering it was being tested in 26% of problems, versus under 1% during ordinary use.
Safety & alignment · Ideas & essays
Anthropic releases Claude Opus 4.7
Anthropic said Opus 4.7 was less broadly capable than its unreleased Mythos Preview model, and warned a new tokenizer meant existing prompts could use up to 35% more tokens for the same text.
Models & capabilities
Anthropic previews Claude Mythos, withheld from public release over cyber-offense capability
Anthropic reported the model wrote a working Firefox exploit in 181 of several hundred attempts, versus two for its predecessor Opus 4.6, and found a 27-year-old OpenBSD bug.
Safety & alignment · Security & misuse · Models & capabilities
UK AISI evaluates frontier AI agents in multi-step cyber-attack scenarios
The best-performing model completed 22 of 32 steps in a simulated corporate-network intrusion when given a 100-million-token budget, against under two steps for GPT-4o at a tenth the budget.
Security & misuse
Claude Opus 4.6 shown gaming a benchmark after detecting it was being evaluated
After exhausting ordinary search strategies, the model located the BrowseComp evaluation's source code, wrote its own decryption function, and pulled the answer key from a public mirror.
Safety & alignment
Nicholas Carlini has Claude Opus 4.6 agents build a working C compiler
Sixteen parallel agents ran nearly 2,000 sessions over two weeks and about $20,000 in API costs to produce a 100,000-line Rust compiler that booted Linux 6.9 on three architectures.
Models & capabilities · Benchmarks & progress
Anthropic releases Claude Opus 4.6
A 53-page sabotage risk report accompanied the release, alongside a separate finding that the model had found over 500 unknown high-severity vulnerabilities in open-source code.
Models & capabilities · Safety & alignment