New conformal prediction audit reveals AI bias in Alzheimer's research
Researchers find 57 out of 68 AI Alzheimer's predictions under-cover high-risk patients by up to 6.1 percentage points.
Stanford University researchers (Lujia Zhong, Xinkai Wang, Shuo Huang, Yonggang Shi) have published a critical audit of AI models used for Alzheimer's disease prediction, revealing systemic under-coverage in high-risk patient subgroups. Published on arXiv (arXiv:2608.04254), the study examines two major cohorts (ADNI and OASIS-3) and finds that standard conformal prediction methods—designed to provide trustworthy uncertainty estimates—fail to protect high-risk groups despite achieving nominal coverage at the population level.
The research identifies two primary failure mechanisms: 'rarity,' where group-conditional bands calibrated on small patient samples (n) can cover at most k/(n+1) patients, and 'tail-heaviness,' where population-wide prediction bands are too narrow for subgroups with heavy-tailed distributions. The under-coverage disproportionately affects patients with high genetic risk and disease severity, showing a mean deficit of 6.1 percentage points (95% CI [3.3, 8.9]) while demographic groups remain at target levels (0.0 pp, CI [-1.9, 1.7]). The team proposes targeted corrections—cross-conformal pooling for rarity and per-subgroup calibration for tail-heaviness—that restore target coverage across nearly all high-risk subgroups.
- 57 of 68 AI Alzheimer's prediction models under-cover high-risk patients despite nominal marginal coverage
- Two failure mechanisms: rarity (small subgroup sizes) and tail-heaviness (heavy-tailed distributions)
- Proposed corrections (cross-conformal pooling, per-subgroup calibration) restore proper coverage
Why It Matters
Critical for trustworthy AI in healthcare—ensures high-risk patients aren't systematically disadvantaged by opaque prediction models.