OpenAI releases MentalHealthBench
On the clinician-authored rubric, GPT-6 Astra scored highest at 57%, ahead of other OpenAI models, Claude Opus 5.5 and Gemini 2.5 Pro, which scored lowest.
- Benchmarks & progress
- Safety & alignment
- Minor
OpenAI released MentalHealthBench, a public benchmark of 1,215 synthetic mental-health conversations paired with 5,262 rubric criteria, which the company said it built with more than 80 licensed psychologists and psychiatrists across 22 countries and 19 languages. Conversations were split across non-acute, high-acuity and emergency situations, and covered adults, teenagers, clinicians and caregivers as user types.
OpenAI reported that among the models it tested, GPT-6 Astra scored highest on the task-clipped rubric score at roughly 57%, ahead of other OpenAI models it evaluated, with Claude Opus 5.5 close behind at around 52%; Gemini 2.5 Pro scored lowest among the frontier models tested, at roughly 30%, similar to GPT-4o’s score from March 2025. For reference points, OpenAI reported that rubric-aware completions written to the criteria scored around 99%, while independent expert-authored completions scored around 39% — a gap the company used to illustrate how much headroom the rubric was designed to leave even for careful human answers.
The release followed a run of scrutiny over chatbot behaviour in mental-health crises, including Transluce’s independent, cross-developer evaluation published weeks earlier, which had tested 77 model variants from six developers rather than OpenAI’s own models alone. Where Transluce’s contribution was independent, comparative measurement, MentalHealthBench is a benchmark OpenAI built, publishes and can score its own models against — useful as a structured, clinician-grounded rubric, but not a substitute for evaluation by a party without a stake in the results.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 24 September 2026 · Shakeel Hashim · TransformerHacking is the least worrying part of OpenAI’s Australia incident
- 25 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseOn Ezra Klein’s Podcast With Jensen Huang
- 12 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseGPT-6-Astra Can Do Ambitious Things