A New Standard for Mental Health Evaluation
Most existing AI evaluations in the mental health domain focus exclusively on emergency crisis detection. MentalHealthBench addresses the gap in assessing how models perform across the full continuum of mental health—from everyday stress and relationship challenges to high-acuity distress. The benchmark uses synthetic, privacy-preserving conversation scenarios to test model responses across diverse personas, including adults, teens, and caregivers, in multiple languages and regions.
Expert-Driven Rubrics and Methodology
OpenAI collaborated with more than 80 licensed psychologists and psychiatrists across 22 countries to build the evaluation framework. The methodology involves:
- Expert-Authored Rubrics: Experts reviewed synthetic conversations and created criteria for ideal model responses. Each criterion is assigned a weight between -10 and +10, where positive values reward beneficial behaviors (e.g., asking clarifying questions) and negative values penalize harmful ones.
- Consensus-Based Scoring: Each conversation was reviewed by at least three experts; only criteria agreed upon by at least two experts (and not contradicted by a third) were retained.
- Automated Grading: The benchmark uses an automated grader (GPT-5.6 Sol) to evaluate model responses against these expert-defined rubrics.
Bridging Clinical Guidance and User Experience
OpenAI conducted a parallel study with 44 users who had previous experience using AI for emotional support to compare expert-defined criteria with user preferences. The findings revealed that while experts prioritize gathering context and clinical nuance, users place a higher premium on practical next steps and the tone of the response. This research highlights that while expert-informed benchmarks are essential for safety, they represent only one dimension of what makes an AI interaction feel helpful and supportive to a human user.