Research & Papers

Martingale-Based Test Confirms Label-Shift Corrections On-the-Fly

Turns routine NLPD monitoring into a formal anytime-valid sequential test.

Deep Dive

In small-batch scientific deployments, labeled target outcomes may be too scarce for reliable shift estimation, even when unlabeled target inputs are available. Seungjin Choi's paper addresses the complementary setting where a practitioner has a pre-specified label-shift correction from domain knowledge and needs to know whether incoming labeled outcomes support it. The key insight is that the per-observation likelihood ratio between a label-shift-corrected predictive model and the source predictive model is a conditional e-value. Consequently, the running product of these ratios is a nonnegative martingale, and Ville's inequality yields an anytime-valid confirmation rule. The log martingale equals the cumulative negative log-predictive density (NLPD) gap between the source and the corrected predictive, turning routine model monitoring into a formal sequential test.

Importantly, rejection of the null hypothesis means the incoming data support the posited correction relative to the source predictive, but it is not a precise estimate of the shift magnitude. Closed-form solutions are available for Gaussian process sources with Gaussian label-shift ratios. Simulations with GP regression validate Type I error control, finite-sample power, sensitivity to miscalibration, and the small-batch advantage of using a reliable prior over label-based re-estimation. The method is particularly valuable when labeled target outcomes are scarce, as it avoids the need for explicit shift re-estimation and instead leverages domain knowledge.

Key Points
  • Per-observation likelihood ratio acts as a conditional e-value, enabling a martingale-based test without requiring a fixed sample size.
  • Cumulative NLPD gap between source and corrected predictive serves as the test statistic, bridging model monitoring and hypothesis testing.
  • Closed-form solutions exist for Gaussian process sources; simulations confirm Type I error control, power, and robustness to miscalibration.

Why It Matters

Unobtrusive validation of label-shift corrections in low-data scientific deployments, saving retraining costs.

📬 Get the top 10 AI stories daily