89 geospatial AI models evaluated: a third unusable beyond source code
New study rates 89 GeoFMs on usability—most fail real-world ecologists.
A new arXiv paper from Robin Young and colleagues at Cambridge systematically evaluates 89 geospatial foundation models (GeoFMs) from a usability perspective. While most prior work focuses on model architecture and benchmark accuracy, this study centers on whether ecologists and conservation scientists can actually adopt these systems. Through a pilot expert elicitation survey, the authors identified misalignments between GeoFM development priorities and the needs of the intended user community. They then built a seven-dimension evaluation framework grounded in HCI theory: Access & Deployment, Interaction & Customization, Trust & Transparency, Community & Support, Scientific Permanence, Multilingual Support, and Offline Usability.
Two raters applied this rubric to 89 GeoFMs, revealing stark accessibility gaps—nearly a third provide no support to practitioners beyond raw source code. Models like Prithvi, Clay, and DynamicWorld scored relatively well on access and deployment but trailed on multilingual and offline usability. The researchers found that dimensions with high rater consistency, such as Community & Support, serve as field-level diagnostics, highlighting where the entire GeoFM ecosystem needs improvement. The paper concludes that model-centric evaluation hides critical usability barriers, and urges developers to prioritize user-centered design to unlock GeoFMs' transformative potential in environmental monitoring.
- Evaluated 89 GeoFMs using a new 7-dimension usability rubric grounded in HCI theory
- Nearly a third (≈29%) of models offer no support beyond source code for practitioners
- High-consistency dimensions like Community & Support act as diagnostics for field-level gaps
Why It Matters
For ecologists, this means many GeoFMs are inaccessible in practice, slowing real-world environmental monitoring adoption.