Research & Papers

Study: Classical EEG features beat foundation models on clinical benchmarks

Five EEG foundation models decoded dataset identity perfectly — exposing a major benchmark flaw.

Deep Dive

A new arXiv preprint from Marzieh Zare (arXiv:2607.24519) delivers a sobering reality check for EEG foundation models in clinical settings. The study evaluated five models — including REVE, CBraMod, and BIOT-bipolar16 — across five tasks and four benchmark datasets plus the Korean CAUEEG cohort, using subject-disjoint validation. On CAUEEG normal/mild cognitive impairment/dementia classification (1,187 recordings), classical features achieved a macro-AUROC of 0.734, beating BIOT-bipolar16 (0.677), CBraMod (0.669), and REVE (0.568). Even a fully randomly initialized encoder scored higher than pretrained REVE (0.667 vs. 0.570), suggesting pretraining does not guarantee clinical utility.

The most striking finding: every encoder could decode dataset identity with 1.000 AUROC, both before and after in-fold PCA-50, and even with balanced subsamples. This proves the models are learning dataset-specific artifacts — montage, cohort, or probe designs — rather than generalizable neural signals. These results held even when label permutations collapsed to chance, confirming the effect is about dataset membership, not causal brain activity. On CHB-MIT cross-subject ictal detection, REVE reached 0.793 AUROC, only marginally better than a nonlinear comparator (0.739); the 5.38-point difference carried a 95% CI that included zero, leaving superiority unresolved. The paper distills these findings into a concrete reporting protocol for clinical EEG foundation-model studies: montage matching, patient-overlap checks, stronger comparators, and representation controls that expose dataset leakage.

Key Points
  • Classical features hit 0.734 macro-AUROC on CAUEEG, beating BIOT-bipolar16 (0.677), CBraMod (0.669), REVE (0.568)
  • All five foundation models decoded dataset identity at 1.000 AUROC, proving they rely on dataset artifacts
  • Random initialization outperformed pretrained REVE on CAUEEG (0.667 vs 0.570), challenging foundation-model value

Why It Matters

EEG foundation-model hype needs a reality check: without negative controls, clinical AI may learn dataset identity, not patient pathology.

📬 Get the top 10 AI stories daily