Researchers launch OVEarth-Bench for AI Earth observation
New benchmark evaluates how well AI handles real-world satellite queries
A team of seven researchers from institutions including Peking University and Nanjing University has released OVEarth-Bench, a new benchmark designed to evaluate AI models in open-vocabulary Earth observation (EO). Unlike existing narrow benchmarks, OVEarth-Bench introduces two key dimensions: category breadth—covering hierarchical geospatial concepts with positive/negative expressions—and query diversity, which includes vocabulary, referring, and reasoning queries. The benchmark supports both mask and box localization under a zero-shot protocol.
The evaluation of 15+ general and EO-specific methods reveals critical gaps: current approaches struggle with broad category coverage, with model rankings stabilizing only under wider vocabularies. Multi-modal large language models (MLLMs) like [unspecified models] currently lead in performance, while EO-specific methods generally underperform general models and rarely match top performers. The findings underscore the need for more realistic, diverse, and large-scale benchmarks to drive reliable progress in open-vocabulary EO. The dataset and evaluation tools are publicly available at the provided URL.
- OVEarth-Bench tests AI on hierarchical geospatial concepts with positive/negative expressions and three query types (vocabulary, referring, reasoning)
- Evaluation of 15+ models shows MLLM-based methods lead, while EO-specific models lag behind despite being tailored for satellite data
- Broader category coverage stabilizes model rankings, highlighting the need for more diverse and realistic benchmarks
Why It Matters
Accelerates development of AI systems capable of accurately interpreting natural language queries for satellite imagery