LOGOS detects objects in aerial imagery with text prompts
New transformer model LOGOS uses text prompts to find objects in satellite images with 20% higher accuracy
Researchers Trong-Thuan Nguyen and Minh-Triet Tran (University of Science, VNU-HCM) have developed LOGOS (Language-guided Oriented Object Detection in Aerial Scenes), a novel transformer-based approach that uses textual prompts to detect oriented objects in aerial imagery. The model addresses key challenges in remote sensing like angular discontinuity and cluttered backgrounds by incorporating prompt-modulated content queries that dynamically adjust focus based on text inputs.
In evaluations on the DOTA dataset, LOGOS achieved state-of-the-art performance, particularly excelling in densely packed and rotated object scenarios. The approach represents a significant advancement over traditional methods by combining language guidance with transformer architectures to improve both robustness and scalability in oriented object detection for geospatial applications.
- LOGOS is a transformer-based model using text prompts for aerial object detection
- Outperforms existing methods by 20% on DOTA dataset for rotated/dense objects
- Handles angular discontinuity and cluttered backgrounds better than traditional approaches
Why It Matters
Transforms satellite imagery analysis by enabling precise object detection through natural language queries, critical for urban planning and disaster response.