Research & Papers

Perets & Mannor's new FDR framework handles finite data and structured hypotheses

New method controls false discovery rate even with limited null samples and complex hypothesis spaces.

Deep Dive

Binyamin Perets and Shie Mannor address a critical problem in scientific discovery: false discovery rate (FDR) control when resources are finite. Traditional FDR methods assume the null distribution is perfectly known or has infinite samples, but real-world experiments often yield only a limited number of null draws, leaving p-values uncertain. Meanwhile, hypothesis spaces in many fields (genomics, neuroscience, etc.) have inherent structure (e.g., clusters, graphs) that should be exploited.

The authors present a framework that handles both finite-data uncertainty and structured hypothesis spaces by representing structure through a reproducing kernel Hilbert space (RKHS). This RKHS framework allows the derivation of two decision rules. The first guarantees exact FDR control (conservative but strict). The second maximizes statistical power by adapting mirror-statistic control into count space, with an analytical assessment of FDR control even when exact mirror symmetry is relaxed.

The tractability provided by the RKHS also enables direct investigation of finite-data uncertainties. The researchers leverage this to suggest a policy for efficient allocation of null distribution samples — a practical guide for researchers deciding how many null draws to generate for each hypothesis. This policy could significantly reduce computational or experimental costs without sacrificing FDR control. The work is particularly relevant for settings like genome-wide association studies or A/B testing with limited control data, where both structure and data scarcity matter.

Key Points
  • Two decision rules: one guarantees exact FDR control, another maximizes power using an adapted mirror statistic in count space.
  • Uses RKHS to represent arbitrary structure in hypothesis spaces, making it applicable to graphs, clusters, and other structured domains.
  • Provides a policy for efficient allocation of finite null distribution samples, reducing resource requirements in large-scale testing.

Why It Matters

Enables more reliable discoveries in data-scarce domains like genomics and A/B testing, saving resources while controlling errors.

📬 Get the top 10 AI stories daily