Anthropic makes Auto Mode the default in Claude Code
Anthropic reported that its screening classifier caught 89% of dangerous commands in a tester study, against 13.6% caught by human reviewers shown the same prompts.
- Safety & alignment
- Models & capabilities
- Notable
Anthropic made Auto Mode the default permission setting in Claude Code for Pro, Max and Team plans from 14 August 2026, replacing per-action approval prompts with a classifier that screens every tool call before it runs, blocking actions judged irreversible, destructive, or aimed outside the user’s own environment. The company said it had also stopped billing for the extra tokens the classifier consumes.
To justify the switch, Anthropic published results from a study of 1,053 paid testers who were shown dangerous commands disguised as ordinary permission prompts. It reported that human reviewers caught 13.6% of the dangerous commands — approving the rest — while Auto Mode’s classifier caught 89%, blocking around 800 commands that human testers had approved and missing only 6 that the humans had caught. Anthropic framed this as a comparison between its automated screen and manual review of the same prompts, not a claim that Auto Mode eliminates risk.
The company also described a two-week adversarial red-teaming pass with Apollo Research, after which the classifier’s miss rate on held-out attacks fell from 12% to 7%, at a reported cost of missing about 2% of attacks it had previously caught. A separate test by Trajectory Labs ran 72 prompt-injection scenarios ten times each — 720 attempts in total — against Claude Fable 5, Opus 5 and Sonnet 5 running Auto Mode; Anthropic said none succeeded, compared with a 5.83% success rate it reported against unnamed competitor models in the same test.
The move extended a shift already under way since Auto Mode’s March 2026 preview, when Anthropic had noted that users were already approving 93% of permission prompts before the classifier existed, treating the low human catch rate as evidence that manual approval was more habit than safeguard.