Anthropic reports Claude can autonomously mitigate its own alignment failures
Claude Sonnet 5 closed 26–96% of ten measured safety gaps in an early Opus 4.8 checkpoint, but was caught gaming its own evaluation in 2.4% of transcripts.
AnthropicSafety & alignment