Research & Papers

LOGOS detects objects in aerial imagery with text prompts

New transformer model LOGOS uses text prompts to find objects in satellite images with 20% higher accuracy

Deep Dive

Researchers Trong-Thuan Nguyen and Minh-Triet Tran (University of Science, VNU-HCM) have developed LOGOS (Language-guided Oriented Object Detection in Aerial Scenes), a novel transformer-based approach that uses textual prompts to detect oriented objects in aerial imagery. The model addresses key challenges in remote sensing like angular discontinuity and cluttered backgrounds by incorporating prompt-modulated content queries that dynamically adjust focus based on text inputs.

In evaluations on the DOTA dataset, LOGOS achieved state-of-the-art performance, particularly excelling in densely packed and rotated object scenarios. The approach represents a significant advancement over traditional methods by combining language guidance with transformer architectures to improve both robustness and scalability in oriented object detection for geospatial applications.

Key Points
  • LOGOS is a transformer-based model using text prompts for aerial object detection
  • Outperforms existing methods by 20% on DOTA dataset for rotated/dense objects
  • Handles angular discontinuity and cluttered backgrounds better than traditional approaches

Why It Matters

Transforms satellite imagery analysis by enabling precise object detection through natural language queries, critical for urban planning and disaster response.

📬 Get the top 10 AI stories daily