arXiv Computer Vision

Infra-Bench CLS: A Global, Open-Source Benchmark for Critical Infrastructure Classification with Earth Observation Foundation Models

arXiv Computer Vision
Sep 14

No One Knows the State of the Art in Geospatial Foundation Models

The paper "No One Knows the State of the Art in Geospatial Foundation Models" critiques the current lack of standardization in geospatial foundation model (GFM) research, highlighting inconsistencies in evaluation, training, and model release practices across 152 papers. It reports significant discrepancies—46 cross-paper disagreements of at least 10 points for the same model and benchmark, 94 out of 126 papers using unique pretraining configurations, and 39% of papers releasing no model weights. The authors propose six concrete expectations, including named-license weight release, shared core evaluations, and a unified evaluation harness, to address these coordination failures and foster a clearer, comparable understanding of GFM progress.

By Isaac Corley, Nils Lehmann, Caleb Robinson, Gabriel Tseng, Anthony Fuller, Hamed Alemohammad, Evan Shelhamer, Jennifer Marcus, Hannah Kerner
arXiv AI
Jul 22

Now We Know? A Systematic Comparison of TerraMind and THOR

arXiv:2607. 18504v1 Announce Type: cross Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact?

By Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-B{\o}rre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci
arXiv Machine Learning
Aug 4

Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

arXiv:2608. 00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.

By Fengxiang Wang, Qiuyang Yu, Yueying Li, Mingshuo Chen, Chengchi Fei, Kaiyi Xu, Lixin Gu, Wangxu Wei, Junchao Gong, Lipeng Ma, Jiong Wang, Fenghua Ling, Wenlong Zhang, Xue Yang, Wenjing Yang, Ben Fei, Long Lan
arXiv Computer Vision
3d ago

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Mapping

Cryo-Bench is a new benchmark that evaluates foundation models for cryosphere mapping, comprising six semantic‑segmentation datasets across five cryospheric components (supraglacial debris, glacial lakes, sea ice, calving fronts, and Antarctic ice‑shelf extent). The benchmark includes multispectral, RGB, and SAR observations from under‑represented regions and tests thirteen geo‑foundation models alongside U‑Net and Vision Transformer baselines. Results show that with frozen encoders U‑Net slightly outperforms TerraMind, but the difference is not statistically significant; fine‑tuning with learning‑rate optimization can dramatically improve performance for some models, while in few‑shot scenarios several foundation models retain over 90 % of their full‑label accuracy.

By Saurabh Kaushik, Lalit Maurya, Beth Tellman, Swalpa Kumar Roy, Valerio Marsocci, Gustau Camps-Valls, Jocelyn Chanussot
arXiv Machine Learning
Aug 19

Spatially explicit feature importance for building height estimation using research-access high-resolution SAR and optical sensors

The study presents a method for estimating building heights in a large Brazilian city using freely available satellite data, including TerraSAR-X StripMap, PlanetScope, and Sentinel-1. A geographically weighted random forest model achieved an RMSE of 5.34 m and an R² of 0.756 against LiDAR reference data, with local feature importance varying by building type and context. The results highlight that no single sensor dominates across all scenarios, offering guidance for selecting satellite-derived products in different urban settings.

By Guilherme Iablonovski, Pierre-Louis Frison, Tatiana Silva da Silva
arXiv Machine Learning
Aug 6

Benchmarking Deep Learning Models for Dense Event Classification of Offshore Wind Infrastructure in Sentinel-1 Time Series

arXiv:2608. 04706v1 Announce Type: new Abstract: Monitoring of offshore wind energy infrastructure life cycles, especially during the deployment phase, is an important contribution for stakeholders to make informed decisions in a phase of increasing deployment activities.

By Thorsten Hoeser, Felix Bachofer, Claudia Kuenzer
arXiv Computer Vision
Sep 3

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

The paper presents a lightweight method to adapt general‑purpose vision‑language models (VLMs) for multispectral and synthetic aperture radar (SAR) image understanding. By rendering each observation as five optical views and one SAR view, naming them in the prompt, and applying LoRA to the language network and selected visual transformer blocks, the authors enable VLMs to process band composites, spectral indices, and radar backscatter without retraining a new foundation model. On a balanced six‑class land‑cover benchmark from BigEarthNet‑v2, the adapted Qwen3‑VL achieves a micro F1 of 0.8275, and the same protocol improves four other VLMs and transfers to flood verification and captioning tasks. "whyItMatters":"The study shows that existing VLMs can be repurposed for multispectral and SAR tasks through simple input rendering and compact LoRA adaptation, avoiding the need for dedicated encoders and domain pretraining."

By Shanji Liu, Kelu Yao, Junxiao Xue, Chenghui Lv, Xiangyang Miao, Yekai Huang, Yaying Chen, Chao Li