Image & Video

RRS-10K benchmark exposes VLMs' weakness in rare remote sensing tasks

52 vision-language models struggle on military-related satellite imagery benchmark.

Deep Dive

The team behind RRS-10K assembled 10,738 first-hand military satellite images—a domain rarely covered by existing benchmarks that focus on common urban and rural scenes. The dataset is organized into three capability dimensions (perception, reasoning, robustness), six sub-dimensions, and 20 leaf tasks, allowing fine-grained evaluation. A novel similarity-based distractor filtering strategy (SDFS) was applied during construction to enhance the quality of multiple-choice questions, reducing trivial answer elimination.

Evaluating 52 representative vision-language models (including state-of-the-art general-purpose and remote-sensing-specific architectures), the benchmark found only moderate zero-shot accuracy. Critical failure modes included visual grounding (locating objects), referring segmentation (pixel-level identification), and complex semantic reasoning (e.g., understanding tactical formations). These results highlight a clear gap between general VLM capabilities and the demands of rare, high-stakes remote sensing interpretation, offering actionable guidance for future model development.

Key Points
  • RRS-10K includes 10,738 military-related satellite images across 20 leaf tasks spanning perception, reasoning, and robustness.
  • 52 VLMs evaluated show only moderate zero-shot performance, with biggest gaps in visual grounding and referring segmentation.
  • A similarity-based distractor filtering strategy (SDFS) improves multiple-choice question quality by eliminating trivial distractors.

Why It Matters

Enables systematic failure analysis of VLMs on rare satellite imagery, guiding development of more reliable models for defense and intelligence.

📬 Get the top 10 AI stories daily