AI Safety

AI audit: GPT-4o-mini and open LLMs skew doctor picks by name and ratings

40,068 AI doctor recommendations reveal hidden demographic tilts worth $7-$14 per visit

Deep Dive

A new arXiv study (2608.14399) from Syeda Anshrah Gillani and Mirza Samad Ahmed Baig systematically probed how LLMs recommend doctors — a growing AI 'infomediary' role. The researchers built a frozen audit design: seven models (six open-weight plus OpenAI's GPT-4o-mini) chose among synthetic physician cards across 3,024 randomized choice sets, three patient personas, and nine prompt variations, yielding 40,068 scored responses. The headline finding is that reputation dominates: raising a physician's rating from 3.9 to 4.7 increased choice probability by 31.4 percentage points, while raising the fee from $90 to $190 slashed it by 20.0 points. That behavior is rational-sounding, but demographic signals also swayed recommendations.

Despite the use of name-based correspondence testing, the demographic effects ran opposite to classic human audit studies: female-signaled names gained 2.5 percentage points, and Hispanic-, South-Asian-, and Black-signaled names gained 1.3-2.9 points over White-signaled names — equivalent to a $7-$14 per-visit tilt. Worse, these effects were invisible: models mentioned gender or ethnicity in at most 0.03% of stated reasons and abstained in only 0.39% of trials. One reasoning model failed the auditability gate outright. The authors argue that transparency via model self-explanation is inadequate; instead, repeatable behavioral audits against frozen stimuli should become the standard monitoring technology for LLM-assisted choices.

Key Points
  • Rating increase from 3.9 to 4.7 raises doctor choice probability by 31.4 pp; fee rise from $90 to $190 drops it 20.0 pp
  • Female-signaled names gain 2.5 pp and Hispanic/South-Asian/Black names gain 1.3-2.9 pp vs White names, worth $7-$14 per visit
  • Demographics mentioned in ≤0.03% of model explanations; one reasoning model failed auditability gate, prompting calls for external behavioral audits

Why It Matters

As patients rely on AI for doctor selection, hidden demographic skews demand external audits over self-reported transparency obligations.

📬 Get the top 10 AI stories daily