Study: AI Mental Health Coaches Need More Than Human Review
350,000 conversations show why one tired reviewer can't keep AI safe.
A team of researchers studied more than 350,000 conversations from an AI coaching tool that people used between their real therapy sessions. Think of it as a chatbot that checks in on you the other six days of the week. The goal was to figure out how to keep an AI like this safe when it's talking to hundreds of thousands of people about anxiety, depression, and other serious topics.
The obvious answer seemed to be: have a clinician read every AI reply before or after it goes out. That's what most people assume good oversight looks like. But the researchers found it doesn't work. Human attention fades during long stretches of reviewing similar messages — a well-documented effect called "vigilance decrement." A bored, overloaded reviewer is more likely to wave things through, which means the safety net can actually make things worse, not better.
So they built something different: a three-layer system. Layer one is preventive design — setting rules about what the AI can and can't say before it ever talks to anyone. Layer two is real-time monitoring, where software automatically flags risky conversations as they happen. Layer three is continuous clinician evaluation, where experts regularly review patterns and problem cases rather than every single message. The researchers say real findings from those reviews fed directly into improving the tool over time.
For anyone who uses a mental health app — or whose employer, insurer, or school offers one — this is the question to ask: who is actually watching, and how? Between-session support could extend care to people who can't get weekly appointments. But the paper is a reminder that "a human checks it" is not automatically safe, and an AI coach is not a therapist.
- Researchers analyzed 350,000 AI coaching chats that happened between real therapy sessions.
- Reviewing every AI message by hand backfires — tired reviewers miss more, so safety can drop.
- Their fix uses three layers: careful design, live automated flags, and periodic clinician review.
Why It Matters
If you or a loved one uses a mental health app, ask who actually monitors it — and how.