Image & Video

CardioState-JEPA unifies ECG, PPG, PCG, boosting murmur detection by 18.8 AUROC

A single cardiac foundation model learns shared physiology across three signal modalities, beating unimodal baselines.

Deep Dive

CardioState-JEPA is a new cardiac foundation model that breaks the single-modality mold by learning one shared representation across electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG). Built by Hamza Shafiq, Aaqib Saeed and colleagues, the model uses a physiology-aware joint-embedding predictive architecture (JEPA) to map heterogeneous waveforms into a common token space, process them with a single Transformer encoder, and pretrain by predicting masked latent cardiac states. Crucially, the pretraining target is shared cardiac physiology, not sensor-specific waveform appearance, so the model learns what ECGs, PPGs, and PCGs have in common rather than memorizing each signal's quirks.

The model tackles the temporal offset between electrical, mechanical, and hemodynamic events using a learned delay aligner that synchronizes cross-modal predictions in cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data, then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks, it improved average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised single-signal baseline. It also matched or exceeded cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks, suggesting heterogeneous cardiac signals can mutually supervise a single foundation model without needing expensive labels.

Key Points
  • CardioState-JEPA uses a joint-embedding predictive architecture (JEPA) with a shared Transformer encoder to learn one representation from ECG, PPG, and PCG signals.
  • A learned delay aligner matches electrical, mechanical, and hemodynamic events, handling temporal offsets between sensor modalities.
  • On 25 downstream tasks, it beats self-supervised baselines by 8.2 AUROC (PPG), 18.8 AUROC (PCG murmur), and 15.5 AUROC (ECG), rivaling clinically supervised models.

Why It Matters

One cardiac foundation model can replace three unimodal systems, enabling more robust, label-efficient diagnostics across ECG, PPG, and PCG.

📬 Get the top 10 AI stories daily