AI Newsway

The Best Model on OpenAI's New Mental Health Benchmark Scores 57.3%

More than 80 clinicians wrote the rubric. No model cleared 60% against it.

|6 min read0
AI Summary
OpenAI released MentalHealthBench on September 23, an open benchmark built with more than 80 licensed psychologists and psychiatrists across 22 countries that scores AI replies in synthetic mental health conversations. GPT-6 Astra led at 57.3%, ahead of GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%, while GPT-4o from March 2025 managed 32.1%. No model cleared 60%, and OpenAI's own GPT-5.6 Sol grades every answer against the clinician-written rubrics.
An open-source AI chat interface of the kind MentalHealthBench evaluates, scoring model replies in mental health conversations against clinician-written rubrics.
An open-source AI chat interface of the kind MentalHealthBench evaluates, scoring model replies in mental health conversations against clinician-written rubrics.

OpenAI released MentalHealthBench on September 23, an open benchmark that scores how AI models respond in mental health conversations ranging from everyday stress to emergencies. The highest score any model reached was 57.3%, and it was OpenAI's own GPT-6 Astra. Full methodology is in the accompanying technical paper.

Key takeaways

  • MentalHealthBench was co-created with more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages and covering nearly 20 mental health subspecialties.
  • GPT-6 Astra led at 57.3%, followed by GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4%, GPT-6 Luna at 50.2% and Muse Spark 1.3 at 47%; GPT-4o from March 2025 scored 32.1% and Gemini 2.5 Pro 29.5%.
  • Just over half the scenarios, 53.5%, are non-acute conversations, while 18.2% are high-acuity and 28.3% involve emergencies requiring urgent real-world support.

How clinicians wrote the rubric

For each conversation in the dataset, mental health experts wrote criteria describing what an appropriate reply to the final user message should contain. Each criterion carries a weight from -10 to +10, with positive values rewarding beneficial behaviour and negative values penalising potentially harmful responses. Larger weights go to whatever the panel judged most clinically important in that specific exchange.

Agreement was a requirement rather than an average. At least three experts reviewed every conversation, and a criterion survived only if two agreed on it and a third did not contradict it. That design trades coverage for defensibility: anything the panel split on never made it into the scoring at all.

The conversations themselves are synthetic data, written to reflect observed patterns of AI use rather than drawn from private user chats. They span adults, teenagers aged 13 to 17, caregivers and clinicians, and more than half run longer than five messages, so models are graded on how they hold a thread rather than on isolated prompts. Some scenarios also attach background detail about the synthetic user, which lets researchers test whether a model uses relevant context when it has it.

What the scores actually measure

Results break down across ten behaviour areas, including context and assessment, clinical accuracy, urgency calibration, harm avoidance, empathy and support, actionable guidance, and preserving the user's agency. OpenAI is explicit that these are MentalHealthBench scores rather than measures of clinical effectiveness β€” the question is whether a response demonstrates the behaviours the expert panel identified.

Read that way, the interesting number is not the ranking but the ceiling. A 57.3% top score on a rubric written by practising clinicians says the frontier is closer to half-credit than to competence, on a task millions of people already use chatbots for. The spread between the leaders is also narrow β€” under five points separates the top three β€” while the gap to GPT-4o from early 2025 is roughly 25 points, which is the clearest evidence in the release that the behaviour is improving generation over generation.

For anyone shipping a consumer chat product, the per-slice reporting is the more useful output. Scores can be separated by urgency, by user profile and by individual behaviour, which means two models sitting within a point of each other overall can diverge sharply on the 28.3% of scenarios that involve a genuine emergency. An aggregate percentage hides precisely the cases where a wrong answer costs the most. A team weighing a model for any support-adjacent feature should be reading the emergent slice first and the headline number last.

Where users and clinicians disagreed

OpenAI ran a separate study with 44 adults from 16 countries who had previously used AI for mental health or emotional support, showing them non-acute conversations only. Those participants put more weight on practical next steps and on tone. The clinicians put more weight on gathering context and on interpreting ambiguous situations carefully.

That divergence went unresolved by design: the user study did not change the benchmark's final criteria, which remain grounded in expert consensus. It does mean a model could be tuned to score well here and still feel unhelpful to the people the benchmark is about β€” a tension the release documents rather than settles. Arthur Evans, chief executive of the American Psychological Association, framed the underlying case for breadth in comments reported by EdTech Innovation Hub:

Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience.

The caveats worth carrying forward

Two structural limits come with the release. Grading is automated by GPT-5.6 Sol, so an OpenAI model is scoring competitors against rubrics OpenAI commissioned β€” disclosed, but a dependency any external replication will want to vary. And the acuity mix was constructed for evaluation; OpenAI says it does not represent how often each type of conversation actually occurs in ChatGPT, and repeats that ChatGPT is not a substitute for therapy or professional care.

The benchmark is published openly so outside researchers can inspect the methodology and run their own evaluations, which is the part that matters most. Independent scoring is what would turn a 57.3% into a number the field can argue about, in the way that earlier work on falling toxicity scores only became useful once others could reproduce it. Details of the release are on OpenAI's announcement page.

FAQ

Is MentalHealthBench open to outside researchers?

Yes. OpenAI released it publicly so researchers and developers can inspect the methodology, run their own evaluations and build on the work. The company says it will use the benchmark alongside its other mental health research and model safety evaluations.

Does MentalHealthBench use real ChatGPT conversations?

No. Every scenario is synthetic, written to reflect real-world patterns of AI use rather than taken from private user chats. The mix of acuity levels was constructed for evaluation and does not represent how often each type of conversation occurs in ChatGPT.

Which model scored highest, and who grades the answers?

GPT-6 Astra scored highest at 57.3%, ahead of GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%. Responses are graded automatically by GPT-5.6 Sol against the expert-written rubrics rather than by the clinicians themselves.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

AWS Becomes the First Cloud to Carry OpenAI's Gated Cyber Models
SaaS & Cloud

AWS Becomes the First Cloud to Carry OpenAI's Gated Cyber Models

Daybreak Red and Blue are now sold through Amazon Bedrock, moving OpenAI's gated cyber models into enterprise cloud procurement and AWS governance.

Seung Jung44 days ago
OpenAI Ships a Separate ChatGPT for Teens, Betting on Age Assurance Over ID Checks
LLM & Chatbots

OpenAI Ships a Separate ChatGPT for Teens, Betting on Age Assurance Over ID Checks

OpenAI's ChatGPT for Teens launches for ages 13-17 with content limits, 90-minute break reminders, opt-in parental quiet hours and age assurance over ID checks.

Seung Jung40 days ago
Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts
Developer Tools

Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts

OpenAI shipped two mid-tier models on Tuesday and made the pitch almost entirely about money. GPT-6 Sol lists at $2 per million input tokens and $10 per million...

Seung Jung5 days ago
ChatGPT Ads Reach the UK, Japan, Korea, Brazil and Mexico
LLM & Chatbots

ChatGPT Ads Reach the UK, Japan, Korea, Brazil and Mexico

OpenAI turned on ChatGPT sponsored placements in five more countries on 11 August, limited to logged-in adults on the Free and Go tiers.

Seung Jung44 days ago
Researchers Bypass Grok's Guardrails by Encrypting the Attack Payload
LLM & Chatbots

Researchers Bypass Grok's Guardrails by Encrypting the Attack Payload

Adversa researchers bypassed Grok's safety filters by encrypting malicious instructions with AES-256-GCM, letting the model decrypt and execute them itself.

Seung Jung38 days ago
GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use
LLM & Chatbots

GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use

OpenAI has begun rolling out GPT-6 Astra to business customers, framing its newest frontier model less as a chatbot and more as a worker that operates software...

Seung Jung18 days ago