Research & Papers

Group Gaussian Mirror brings FDR control to sequential and grouped models

New arXiv paper controls false discovery rates for lagged, recurrent, and attention-based feature blocks.

Deep Dive

Featured on arXiv as 2608.00989, this paper by Jiaan Han, Junxiao Chen, and Yanzhe Fu tackles a blind spot in FDR-controlled feature selection. Most methods assume coordinate-wise hypotheses, where each feature maps to a single importance score. But in real-world sequential and grouped models—think time series with lags, recurrent states, or attention-based embeddings—one original feature is represented as a block of sub-features. The authors propose a grouped-feature FDR control framework that works for both grouped linear models and nonlinear neural sequential models.

For grouped linear models, they construct null-symmetric block-level mirror statistics using matrix-valued perturbations, proving FDR control for both low- and high-dimensional settings. For neural sequential models, they combine Permutation SHAP derivatives—model-agnostic block-level importance scores—with a kernel-based dependence measure. This avoids needing to specify the covariate distribution and generalizes across architectures. When block size equals one, the method reduces neatly to existing Gaussian Mirror or Neural Gaussian Mirror methods, confirming its theoretical coherence. Experiments on simulated and real-world data demonstrate reliable FDR control and better statistical power, especially when grouped features are correlated. This gives practitioners a principled way to trust feature selections in complex, high-dimensional sequential models where spurious correlations are common.

Key Points
  • Framework supports both grouped linear models and nonlinear neural sequential models with block-level mirror statistics.
  • Uses Permutation SHAP derivatives paired with kernel-based dependence measures, requiring no covariate distribution specification.
  • Proven FDR control in low- and high-dimensional settings, with improved power on correlated grouped-feature signals in experiments.

Why It Matters

Time series and attention-based models gain reliable feature selection, reducing false positives in high-stakes sequential predictions.

📬 Get the top 10 AI stories daily