Research & Papers

GeoFM study reveals layer-wise strategies for remote sensing transfer

Intermediate model layers outperform final embeddings for downstream tasks...

Deep Dive

A new study from University of Colorado Boulder systematically evaluates how self-supervised geospatial foundation models (GeoFMs) transfer to downstream tasks. The researchers tested 6 GeoFMs spanning joint-embedding, reconstruction, and multimodal pretraining families across classification, regression, and segmentation benchmarks. They found that model rankings change substantially depending on the task and adaptation settings, challenging the notion of a single 'best' foundation model.

Layerwise probing revealed that intermediate transformer blocks often hold more task-relevant information than the final-layer embeddings, and that different GeoFMs exhibit distinct depthwise profiles. In segmentation case studies on the PASTIS and Sen1Floods11 datasets, downstream adaptation choices—such as decoder design and fine-tuning strategy—were as impactful as the choice of GeoFM itself. Standard dense-prediction heads may be poorly aligned with how GeoFMs organize information across depth. CKA analysis showed fine-tuning does not rewrite all layers uniformly; the strongest changes are localized to the first linear layer of the MLP in ViT blocks, explaining why benchmark rankings shift and motivating representation-aware evaluation strategies.

Key Points
  • 6 GeoFMs from joint-embedding, reconstruction, and multimodal families tested across classification, regression, and segmentation.
  • Intermediate transformer blocks consistently outperform final-layer embeddings for task-relevant information.
  • Decoder design and fine-tuning matter as much as model choice; fine-tuning primarily rewrites the first MLP linear layer in ViT blocks.

Why It Matters

Practical guidance for selecting and fine-tuning geospatial AI models—intermediate layers and adaptation strategy often beat picking the 'best' model.

📬 Get the top 10 AI stories daily