CARE-MH framework standardizes evaluation of mental health LLMs
New benchmark reveals reproducibility gaps in AI therapy assessments
As LLMs increasingly provide mental health support, reliable evaluation of safety, empathy, and therapeutic appropriateness becomes critical. However, existing benchmarks suffer from inconsistent evaluation designs and metric definitions, making results difficult to compare or reproduce. The CARE-MH framework addresses this by providing a unified, standardized approach to evaluating mental health LLMs.
Using CARE-MH, the researchers reproduced and analyzed current state-of-the-art benchmarks. Their findings reveal that reproducibility is strongly tied to model stability, and cross-benchmark disagreements primarily arise from differing metric definitions. The paper (39 pages, 22 figures, 22 tables) argues for standardized evaluation configurations and shared metric definitions to advance the field responsibly.
- CARE-MH provides a unified framework for comparable and reproducible evaluation of mental health LLMs
- Reproducibility depends strongly on model stability, not just benchmark design
- Cross-benchmark disagreements stem from inconsistent metric definitions rather than model capabilities
Why It Matters
Standardized mental health LLM evaluation ensures safe, empathetic AI therapy tools are reliable and trustworthy.