Research & Papers

PatientAgentBench: New benchmark evaluates safety of patient-facing AI agents in healthcare

Synthetic patients and LLM juries expose critical gaps in AI clinical reasoning

Deep Dive

As healthcare AI evolves from answering static questions to completing multi-step tasks — scheduling appointments, managing prescriptions, triaging symptoms — benchmarks must measure real clinical safety. Existing tests either focus on static medical knowledge (e.g., exam questions) or on tool-using agents for providers, ignoring the nuanced, multi-turn conversations needed with patients.

PatientAgentBench addresses this gap by generating a synthetic patient chart, a realistic clinical vignette, and a patient agent that interacts with the AI system under evaluation. An LLM-as-a-jury panel scores each interaction against reusable criteria. In testing, even capable models showed clinical shortcomings: omitting crisis resources in emergencies or falsely claiming actions. The framework allows scalable evaluation to identify safety gaps and guide improvements for safer patient-facing agents.

Key Points
  • Generates synthetic patient records, clinical vignettes, and patient agents for realistic multi-turn conversations
  • Uses an LLM-as-a-jury panel to score responses against clinician-vetted, reusable criteria
  • Identified recurring failures: missing crisis resources in emergencies and claiming unexecuted actions, even in top models

Why It Matters

Ensures patient-safe AI agents by rigorously testing clinical reasoning, task execution, and escalation protocols

📬 Get the top 10 AI stories daily