EEG Foundation Models Fall into 'Identity Trap' – FMScope Diagnoses Shortcut Learning
Leading EEG AI models may be mistaking patient identity for clinical biomarkers, study shows.
A new paper from UC San Diego researchers (Lin, Wu, Jung) published on arXiv reveals a critical flaw in EEG foundation models: the 'Identity Trap.' These models, including LaBraM, CBraMod, and REVE, often achieve high accuracy on clinical resting-state EEG by relying on subject-identity features rather than genuine clinical biomarkers. The team proposes FMScope, a frozen-representation protocol that packages five diagnostics: variance decomposition, subject-axis erasure, aperiodic 1/f ablation, layer-wise label probing, and within-subject direction consistency. Across four datasets in a 2x2 layout, they show that subject-identity features dominate representations—13 to 89 times more variance than random baselines in all 12 model-dataset pairs. Fine-tuning amplifies this dominance by +10 to +63 percentage points. Critically, erasing the subject-identity axis improves label decoding where labels vary within subjects, yielding +6 to +12 percentage points in primary cells and up to +27 pp across external cohorts.
The analysis also reveals that aperiodic 1/f activity is a major carrier of subject identity: removing it drops subject-probe accuracy by 9–19 pp for LaBraM and CBraMod. However, REVE saturates subject identity without measurable aperiodic dependence, suggesting alternative subject-specific features. Fine-tuning amplifies label variance only in cells where a literature-established cross-subject marker exists, indicating that genuine clinical biomarkers are present but overshadowed by identity shortcuts. The Identity Trap is a physically-grounded instance of shortcut learning—the preferred cue has a measurable physiological component, and subject-disjoint cross-validation alone cannot rule it out. FMScope separates gains reflecting a biological marker from those reflecting subject identity, providing a crucial diagnostic tool for the medical AI community.
- Subject-identity variance is 13–89x random null across all 12 model-dataset pairs, rising further under fine-tuning.
- Erasing subject identity improves label decoding by +6 to +12 pp where labels vary within subjects.
- Aperiodic 1/f activity is a major carrier of subject identity; removing it drops subject-probe accuracy by 9–19 pp for LaBraM and CBraMod.
Why It Matters
For medical AI, claimed EEG diagnostic accuracy may be inflated by subject-specific shortcuts rather than true clinical biomarkers.