New benchmark reveals LLMs diagnose prematurely in psychiatry
Over 60% under-abstention rate even for top models on clinical notes
A new benchmark called Safe-Psych challenges the assumption that LLMs handle incomplete clinical information well. Developed by a team of researchers, Safe-Psych contains over 1,000 real-world psychiatric clinical notes segmented to simulate incremental evidence disclosure. Psychiatrists labeled each stage with one of three actions: DIAGNOSE, CLARIFY, or ABSTAIN. The goal is to test whether models recognize when they lack sufficient evidence to make a reliable diagnosis.
Evaluating multiple state-of-the-art LLMs, the study finds that capability does not ensure calibration. Under incomplete information, under-abstention rates exceed 60% for most models, meaning they diagnose when they should not. Safety-aware prompting reduces premature commitment but shifts errors toward excessive abstention. Crucially, sequential evaluation shows models frequently diagnose before enough evidence is present, and these premature diagnoses are significantly less accurate than those made at the right time. The findings highlight a critical limitation: LLMs struggle to recognize when clinical evidence is incomplete and additional information is needed.
- Under-abstention rates exceed 60% for most models under incomplete information
- Premature diagnoses are significantly less accurate than on-time diagnoses
- Safety-aware prompting reduces overconfidence but increases excessive abstention
Why It Matters
LLMs risk misdiagnosis in psychiatry when they don't ask for more info first