Research & Papers

Prior laundering: AI overconfidence hidden in legacy data

Undetectable overconfidence from training on archive reconstructions threatens medical and seismic imaging.

Deep Dive

A new paper from Ali Siahkoohi and Sina Alemohammad, titled 'Prior laundering: learned priors with inherited, undetectable overconfidence,' reveals a critical flaw in how AI models for inverse problems are trained. The focus is on learned generative priors used in Bayesian inference for ill-posed problems like seismic and medical imaging. Since ground-truth data is scarce, practitioners often train these priors on archives of legacy reconstructions—a process the authors call 'prior laundering.' The problem is that when the measurements are uninformative (i.e., they don't constrain certain directions in the solution space), the posterior uncertainty collapses to reflect the archive's belief, not any actual data evidence. This inherited overconfidence is undetectable in deployment because any truths that differ only in those blind directions will produce the same data likelihood, and self-consistency checks like simulation-based calibration will pass regardless.

The paper proves mathematically that averaging the legacy posterior over measurements yields the old regularizer advanced by a single expectation-maximization step—improved where data resolves, frozen where it cannot. A single-best archive is even worse, collapsing blind credible intervals to zero width. In experiments, a diffusion prior trained on archive data under-covers the operator's blind subspace compared to a truth-trained control, and a normalizing flow does the same on a nonlinear groundwater operator. The authors recommend that practitioners explicitly report which measurement directions are resolved, separating confidence supported by data from belief inherited through the pipeline. This work has immediate implications for any field relying on learned priors from imperfect data, especially where uncertainty quantification is safety-critical.

Key Points
  • Training generative priors on legacy reconstructions (prior laundering) inherits undetectable overconfidence when measurements are uninformative.
  • The posterior reverts to the archive's belief in blind directions, and standard checks like simulation-based calibration cannot detect the error.
  • Diffusion priors and normalizing flows under-cover blind subspaces; authors recommend reporting which measurement directions are resolved.

Why It Matters

Safety-critical AI for medical and seismic imaging may be overconfident due to data scarcity, hiding model blind spots.

📬 Get the top 10 AI stories daily