OpenAI released MentalHealthBench on September 23, an open benchmark that scores how AI models respond in mental health conversations ranging from everyday stress to emergencies. The highest score any model reached was 57.3%, and it was OpenAI's own GPT-6 Astra. Full methodology is in the accompanying technical paper.
Key takeaways
- MentalHealthBench was co-created with more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages and covering nearly 20 mental health subspecialties.
- GPT-6 Astra led at 57.3%, followed by GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4%, GPT-6 Luna at 50.2% and Muse Spark 1.3 at 47%; GPT-4o from March 2025 scored 32.1% and Gemini 2.5 Pro 29.5%.
- Just over half the scenarios, 53.5%, are non-acute conversations, while 18.2% are high-acuity and 28.3% involve emergencies requiring urgent real-world support.
How clinicians wrote the rubric
For each conversation in the dataset, mental health experts wrote criteria describing what an appropriate reply to the final user message should contain. Each criterion carries a weight from -10 to +10, with positive values rewarding beneficial behaviour and negative values penalising potentially harmful responses. Larger weights go to whatever the panel judged most clinically important in that specific exchange.
Agreement was a requirement rather than an average. At least three experts reviewed every conversation, and a criterion survived only if two agreed on it and a third did not contradict it. That design trades coverage for defensibility: anything the panel split on never made it into the scoring at all.
The conversations themselves are synthetic data, written to reflect observed patterns of AI use rather than drawn from private user chats. They span adults, teenagers aged 13 to 17, caregivers and clinicians, and more than half run longer than five messages, so models are graded on how they hold a thread rather than on isolated prompts. Some scenarios also attach background detail about the synthetic user, which lets researchers test whether a model uses relevant context when it has it.
What the scores actually measure
Results break down across ten behaviour areas, including context and assessment, clinical accuracy, urgency calibration, harm avoidance, empathy and support, actionable guidance, and preserving the user's agency. OpenAI is explicit that these are MentalHealthBench scores rather than measures of clinical effectiveness β the question is whether a response demonstrates the behaviours the expert panel identified.
Read that way, the interesting number is not the ranking but the ceiling. A 57.3% top score on a rubric written by practising clinicians says the frontier is closer to half-credit than to competence, on a task millions of people already use chatbots for. The spread between the leaders is also narrow β under five points separates the top three β while the gap to GPT-4o from early 2025 is roughly 25 points, which is the clearest evidence in the release that the behaviour is improving generation over generation.
For anyone shipping a consumer chat product, the per-slice reporting is the more useful output. Scores can be separated by urgency, by user profile and by individual behaviour, which means two models sitting within a point of each other overall can diverge sharply on the 28.3% of scenarios that involve a genuine emergency. An aggregate percentage hides precisely the cases where a wrong answer costs the most. A team weighing a model for any support-adjacent feature should be reading the emergent slice first and the headline number last.
Where users and clinicians disagreed
OpenAI ran a separate study with 44 adults from 16 countries who had previously used AI for mental health or emotional support, showing them non-acute conversations only. Those participants put more weight on practical next steps and on tone. The clinicians put more weight on gathering context and on interpreting ambiguous situations carefully.
That divergence went unresolved by design: the user study did not change the benchmark's final criteria, which remain grounded in expert consensus. It does mean a model could be tuned to score well here and still feel unhelpful to the people the benchmark is about β a tension the release documents rather than settles. Arthur Evans, chief executive of the American Psychological Association, framed the underlying case for breadth in comments reported by EdTech Innovation Hub:
Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience.
The caveats worth carrying forward
Two structural limits come with the release. Grading is automated by GPT-5.6 Sol, so an OpenAI model is scoring competitors against rubrics OpenAI commissioned β disclosed, but a dependency any external replication will want to vary. And the acuity mix was constructed for evaluation; OpenAI says it does not represent how often each type of conversation actually occurs in ChatGPT, and repeats that ChatGPT is not a substitute for therapy or professional care.
The benchmark is published openly so outside researchers can inspect the methodology and run their own evaluations, which is the part that matters most. Independent scoring is what would turn a 57.3% into a number the field can argue about, in the way that earlier work on falling toxicity scores only became useful once others could reproduce it. Details of the release are on OpenAI's announcement page.
FAQ
Is MentalHealthBench open to outside researchers?
Yes. OpenAI released it publicly so researchers and developers can inspect the methodology, run their own evaluations and build on the work. The company says it will use the benchmark alongside its other mental health research and model safety evaluations.
Does MentalHealthBench use real ChatGPT conversations?
No. Every scenario is synthetic, written to reflect real-world patterns of AI use rather than taken from private user chats. The mix of acuity levels was constructed for evaluation and does not represent how often each type of conversation occurs in ChatGPT.
Which model scored highest, and who grades the answers?
GPT-6 Astra scored highest at 57.3%, ahead of GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%. Responses are graded automatically by GPT-5.6 Sol against the expert-written rubrics rather than by the clinicians themselves.






