Level II-A framework exposes when predictive models fool themselves with past data
New statistical framework uses post-endpoint randomization to test if past-adapted explanations are sufficient.
A new Perspective paper by George Sopasakis (Ximantis AB) and Alexandros Sopasakis (Lund University) tackles a core problem in predictive modeling: a model can fit its data even when its information set is insufficient. Fit alone cannot establish sufficiency. The authors propose Level II-A, a design-based inference framework that tests whether past-adapted information actually explains a committed outcome. The approach commits to a pre-event endpoint, then randomizes the delay to the imperative event. This later-assigned delay acts as a negative-control probe: if only pre-commitment information influenced the endpoint, it cannot systematically order by that delay.
The framework is demonstrated with anticipatory EEG using contingent negative variation, though no human data are analyzed—only synthetic benchmarks. Leakage-safe preprocessing, a frozen label-blind comparator, and retained-sample qualifications carry the exclusion to the confirmatory residual. The synthetic benchmark reports grid-based false-adequacy boundaries of 15 µV·s⁻¹ for assignment isolation and 30 µV·s⁻¹ for the sequential e-value route, in both directions. A non-compensatory rule separates diagnostic failure, selection-limited, opposite-direction, and inconclusive outcomes. The design transfers to any setting where endpoint commitment precedes an exogenous label, turning "the past explains it" from an assumption into a magnitude-qualified, testable claim. The paper spans 90 pages with 6 figures and 15 tables, plus supplementary information and reproducible software.
- Level II-A uses post-endpoint randomization as a negative-control probe to test sufficiency of past-adapted explanations.
- Synthetic EEG benchmark yields false-adequacy boundaries of 15 µV·s⁻¹ for assignment isolation and 30 µV·s⁻¹ for sequential e-value routes.
- No human EEG data analyzed; reproducible software and synthetic data provided for broader application.
Why It Matters
Gives researchers a rigorous, quantitative way to verify whether predictive model explanations are truly sufficient, not just data-fitting.