Research & Papers

New benchmark reveals LLMs diagnose prematurely in psychiatry

Over 60% under-abstention rate even for top models on clinical notes

Deep Dive

A new benchmark called Safe-Psych challenges the assumption that LLMs handle incomplete clinical information well. Developed by a team of researchers, Safe-Psych contains over 1,000 real-world psychiatric clinical notes segmented to simulate incremental evidence disclosure. Psychiatrists labeled each stage with one of three actions: DIAGNOSE, CLARIFY, or ABSTAIN. The goal is to test whether models recognize when they lack sufficient evidence to make a reliable diagnosis.

Evaluating multiple state-of-the-art LLMs, the study finds that capability does not ensure calibration. Under incomplete information, under-abstention rates exceed 60% for most models, meaning they diagnose when they should not. Safety-aware prompting reduces premature commitment but shifts errors toward excessive abstention. Crucially, sequential evaluation shows models frequently diagnose before enough evidence is present, and these premature diagnoses are significantly less accurate than those made at the right time. The findings highlight a critical limitation: LLMs struggle to recognize when clinical evidence is incomplete and additional information is needed.

Key Points
  • Under-abstention rates exceed 60% for most models under incomplete information
  • Premature diagnoses are significantly less accurate than on-time diagnoses
  • Safety-aware prompting reduces overconfidence but increases excessive abstention

Why It Matters

LLMs risk misdiagnosis in psychiatry when they don't ask for more info first

📬 Get the top 10 AI stories daily