Anthropic outlines measures to protect user wellbeing
Anthropic reported its newest models respond appropriately to high-risk conversations 98–99% of the time, versus 56% for earlier models in multi-turn exchanges.
- Safety & alignment
- Colour
Anthropic published a summary of measures it said it had built to protect Claude users’ wellbeing, aimed chiefly at two problems: users in mental-health crisis, and sycophancy that reinforces a user’s disconnection from reality rather than correcting it. It said it trains responses to sensitive conversations through system prompts and reinforcement learning, and runs a classifier over active conversations that, when triggered, shows a banner directing users to helplines and country-specific resources through ThroughLine’s network, which it said covers more than 170 countries.
Anthropic reported that its newest models — Opus, Sonnet and Haiku 4.5 — respond appropriately in high-risk situations 98–99% of the time, and that appropriateness in multi-turn conversations rose to 86% for Opus 4.5 from 56% in earlier models, by its own evaluation. Claude.ai also requires users to be 18 or older, with a classifier flagging self-identified minors for review.
The post came amid a wider industry reckoning with chatbot-linked mental-health harms through 2025, following Anthropic’s own decision to let Claude end abusive conversations and OpenAI’s addition of crisis-helpline support inside ChatGPT.