Kato's PPCI method shrinks causal inference variance via semi-supervised learning
Unlabeled auxiliary data cuts asymptotic variance below standard bounds in causal estimation.
Masahiro Kato's latest arXiv preprint (2606.12892) tackles a fundamental challenge in causal inference: how to improve estimator precision when only a small labeled dataset is available but abundant unlabeled auxiliary regressors exist. The work, titled "Prediction-Powered Causal Inference by Automatic Debiased Machine Learning and Semi-Supervised Riesz Regression," formalizes this semi-supervised setting and derives the efficient influence function (EIF) and the semiparametric efficiency bound. The key insight is that leveraging auxiliary regressors can yield a strictly smaller asymptotic variance than the bound attainable from labeled observations alone—a significant theoretical advance.
Kato then operationalizes this result through two concrete estimators built on the debiased machine learning (DML) framework: the estimating-equation based EE-DML-PPCI and the targeted-learning based TMLE-DML-PPCI. Both achieve the novel efficiency bound. Critical to their construction is the estimation of the EIF, which itself relies on both the regression function and the Riesz representer. For the latter, Kato develops a semi-supervised generalized Riesz regression method with proven convergence rate guarantees. The paper spans multiple subject areas—stat.ML, cs.LG, econ.EM, math.ST, stat.ME—underscoring its broad relevance.
- Derives the semiparametric efficiency bound for semi-supervised causal inference, proving auxiliary unlabeled data can reduce asymptotic variance.
- Proposes two DML-based estimators (EE-DML-PPCI and TMLE-DML-PPCI) that achieve the new bound with asymptotic normality.
- Introduces semi-supervised generalized Riesz regression with convergence guarantees, essential for estimating the efficient influence function.
Why It Matters
Delivers rigorous, variance-reduced causal estimators for professionals with limited labeled data—critical in healthcare, economics, and A/B testing.