Fake Patients Now Test AI Mental Health Screenings — And Find Gaps
AI asked more questions but brushed past suicide warnings. That gap could matter.
Mental health systems are starting to use AI chatbots to handle "psychiatric intake" — the first appointment where someone is asked about their history, symptoms, sleep, and safety. Before trusting these tools with real people, hospitals need a routine way to check whether they meet clinical standards. So researchers built a testing platform called InterviewPlayground, where an AI acts as a simulated patient following a carefully written script.
The results of a small pilot are revealing. Six clinicians spent 25 minutes each comparing notes with a GPT-based AI interviewer. The AI was a sponge: it pulled out 88% of the important clinical details hidden in the patient's story, versus 39% for the humans. But being thorough isn't the same as being careful. The AI made unsupported clinical leaps 57% of the time — inferring conditions nobody had mentioned — compared with 28% for clinicians. And when a safety concern came up, it properly characterized that concern only a third of the time, while clinicians did so about two-thirds of the time.
Why does that matter outside a lab? Intake is exactly where red flags — thoughts of self-harm, medication conflicts, psychosis — are supposed to surface. An AI that gathers facts efficiently but downplays danger could send someone down the wrong care path. The researchers' point isn't that AI intake is doomed. It's that hospitals need a repeatable, low-effort way to measure these tools, including the safety-critical parts, before and after they go live.
The simulated-patient approach is clever because it's reusable. Instead of paying clinicians to role-play the same interview repeatedly, a health system can run hundreds of scripted patients through an AI system overnight and see where it slips. For patients, that means the AI doing your first appointment may soon be graded on whether it notices what you're actually afraid to say.
- Researchers created an AI "practice patient" so hospitals can rehearse and grade chatbots that do mental health first appointments.
- The AI interviewer found 88% of important symptoms versus 39% for human clinicians — but invented unsupported conclusions 57% of the time versus 28%.
- The biggest worry: the AI properly flagged safety concerns only 33% of the time, while clinicians managed 67%.
- The pilot was tiny — just 6 clinicians in 25-minute sessions — so treat the numbers as a signal, not proof.
Why It Matters
If AI handles your first therapy appointment, it may catch more symptoms but miss danger signs a human would catch.