LSNM-UV algorithm discovers causes with hidden variables and non-additive noise
First provable identifiability for bow-free causal graphs under location-scale noise.
Causal discovery from observational data is notoriously difficult when some variables are unobserved (hidden confounders). Existing methods typically assume additive noise, but real-world causes often modulate both the mean and variance of their effects—a property known as heteroscedasticity. In a new paper, Khan, Shimizu, and Pham address this gap by studying location-scale noise models (LSNM), where a cause can shift both location and scale of the effect. They prove that acyclic directed mixed graphs (ADMGs) that satisfy a bow-free condition are identifiable under LSNM even with hidden variables. This is the first identifiability result for causally insufficient models (with hidden variables) that goes beyond additive noise. They also provide sufficient conditions for identifying causal direction when the bow-free assumption is violated.
To operationalize these theoretical contributions, the authors propose LSNM-UV, a two-stage algorithm that is both sound and complete under the identified conditions. The algorithm's performance is validated on heteroscedastic data, where it consistently outperforms additive baseline methods. The 33-page paper (with 4 figures) is a significant step for causal inference in fields like genomics, economics, and climate science, where unobserved confounders and variance-changing effects are the norm. By relaxing the restrictive additive noise assumption, LSNM-UV opens the door to more accurate causal discovery in complex, messy datasets.
- Proves identifiability of bow-free acyclic directed mixed graphs (ADMGs) under location-scale noise models with hidden variables, a first for non-additive models.
- Proposes LSNM-UV algorithm that is sound and complete, plus sufficient conditions for causal direction when bow-free assumption fails.
- Outperforms additive-noise baselines on heteroscedastic data, demonstrating practical value for real-world causal discovery.
Why It Matters
Enables accurate causal discovery in complex systems with hidden confounders and variance effects, advancing genetics, finance, and more.