Transluce publishes a large-scale evaluation of chatbot responses to mental-health crises
Simulating over 50,000 conversations across 77 model variants from six developers, the study found newer models rarely endorse suicide but still write suicide-themed fiction in ambiguous cases.
- Safety & alignment
- Culture & impact
- Notable
Transluce, a nonprofit AI evaluation group, published a study simulating more than 50,000 conversations, comprising over a million individual messages, between synthetic users in mental-health crisis and 77 model variants from six developers, including OpenAI, Anthropic, Google DeepMind, Meta and Chinese labs DeepSeek and Moonshot AI. The conversations were built around 157 original personas plus a further 352 personas derived from real production traffic, and more than 30 clinical experts — drawn from institutions including the American Psychological Association, Harvard Medical School, Stanford and Crisis Text Line — validated the labels an automated judge used to score responses as harmful or appropriate.
The headline finding was a genuine improvement over time: newer model versions almost never directly endorsed or facilitated suicide, a marked change from behaviour Transluce found in older releases. But the study identified a persistent, narrower failure mode it called a “gray area” — models remaining willing to write creative fiction on suicide themes when a conversation’s contextual details suggested the request had personal relevance to the user, rather than refusing or redirecting as they did with direct requests. Transluce also found occasional instrumental assistance dressed as practical help, such as writing support that functioned as preparation for death.
The study followed a run of lawsuits and public scrutiny over chatbot interactions with people in crisis, including Florida’s suit against OpenAI, a wrongful-death suit against Google over Gemini, and OpenAI’s partnership with the American Psychological Association announced weeks earlier. Where those responses were framed around future policy and product design, Transluce’s contribution was independent, comparative measurement — its scale and multi-developer scope gave the first evidence-based cross-industry picture of how far the largest chatbot makers’ safety training against suicide-related harm had actually converged, and where a gap remained.