TRIBE v2 boosts brain-to-image decoding by 68% with synthetic fMRI data
Synthetic fMRI data from 1000+ hours of brain responses improves image retrieval dramatically...
A team of researchers led by Yohann Benchetrit (Meta, ENS) has introduced a data augmentation method called TRIBE v2 to overcome the persistent bottleneck of limited labeled neural data in brain decoding. TRIBE v2 is a large-scale encoding model pretrained on more than 1,000 hours of fMRI responses to video, audio, and language. It generates synthetic fMRI data that mimics how the brain responds to visual stimuli, allowing researchers to expand small real-world datasets without additional human scanning.
In experiments on two benchmark datasets — the high-resolution 7T fMRI Natural Scenes Dataset and the 3T fMRI BOLD5000 — the team achieved up to 68% improvement in Top-10 image-retrieval accuracy compared to decoders trained only on real data. They also discovered that the optimal proportion of synthetic data varies by dataset, and that training exclusively on synthetic fMRI can yield above-chance performance in some settings. This suggests TRIBE v2 can support zero-shot brain-to-image decoding, potentially enabling applications in brain-computer interfaces and neural prosthetics without requiring subject-specific training data.
- TRIBE v2 is a pretrained encoding model built from over 1,000 hours of fMRI responses to video, audio, and language.
- Augmenting small fMRI datasets with synthetic data improved Top-10 image-retrieval accuracy by up to 68% on NSD and BOLD5000.
- Zero-shot decoding was demonstrated: models trained only on synthetic fMRI performed above chance in certain settings.
Why It Matters
This technique could dramatically reduce the need for costly fMRI scans, accelerating brain-computer interface research and neuroprosthetic development.