Image & Video

UCSD & MIT's MI-RAFT aligns retinal images across any modality

New two-stage framework beats modality-dependent methods in retinal image registration

Deep Dive

Researchers from UC San Diego and MIT have introduced a new approach to retinal image registration, a critical step in ophthalmic diagnosis and longitudinal disease monitoring. The paper, titled "Modality-Invariant Coarse-to-Fine Retinal Image Registration," addresses a key limitation of existing methods: most are designed for a single imaging modality or a fixed pair of modalities, making them inflexible in clinical settings where multiple retinal imaging types—such as color fundus, OCT, and fluorescein angiography—are used together. The proposed framework overcomes this by operating in two stages. First, a universal retinal vessel segmentation model guides sparse feature matching to achieve robust coarse global alignment across modalities. Second, a novel modality-invariant optical flow estimation network, called MI-RAFT, refines the alignment through dense local registration, delivering high precision even when image appearances differ significantly between modalities.

Submitted to IEEE Transactions on Image Processing (TIP-40498-2026), the work was authored by Bo Wen, Nehal Nailesh Mehta, Melanie Tran, Dirk-Uwe Bartsch, William Freeman, and Truong Nguyen. Extensive experiments show that the method handles diverse combinations of commonly used retinal imaging modalities, demonstrating strong modality invariance while outperforming state-of-the-art modality-dependent registration techniques. By decoupling registration from specific imaging hardware, the framework allows a single tool to align images from any pair of modalities, simplifying multimodal analysis pipelines and improving longitudinal disease tracking. Clinicians could compare scans across different devices or time points without needing separate algorithms for each modality pairing. The paper is available on arXiv (2608.14829) with code and data links provided, making it a practical resource for the medical imaging community.

Key Points
  • Two-stage framework: sparse vessel-based feature matching for coarse alignment, then MI-RAFT optical flow for dense refinement
  • Modality-invariant by design, handling arbitrary retinal imaging modality combinations without retraining or fixed pairs
  • Outperforms state-of-the-art modality-dependent methods in experiments across diverse retinal imaging datasets

Why It Matters

Enables flexible cross-modal retinal image alignment for better ophthalmic diagnosis and longitudinal disease monitoring.

📬 Get the top 10 AI stories daily