The paper introduces a multimodal foundation model for lunar remote sensing, trained from scratch on SomBench—a dataset of nearly two million co‑registered tile bundles across 11 modalities at 1 m and 100 m resolutions. The model extends the TerraMind masked‑token architecture with lunar‑specific features such as explicit acquisition geometry and joint training of two spatial scales, and employs FlexiViT patch embeddings for adaptable patch sizes. Evaluation on crater detection, irregular mare patch segmentation, and polar ice prospectivity regression shows that the pretrained model matches or surpasses ImageNet‑pretrained baselines, with notable label efficiency and effective adaptation via LoRA.
By Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal, Rachel A. Slank, Geoffrey Dawson, Hiyam Debary, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Nikolaos Dionelis, Ankur Kumar, Campbell D. Watson, Manil Maskey, Rebekah I. Dawson-Rigas, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
arXiv:2504. 11171v5 Announce Type: replace-cross Abstract: We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO).
By Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Long\'ep\'e
Cryo-Bench is a new benchmark that evaluates foundation models for cryosphere mapping, comprising six semantic‑segmentation datasets across five cryospheric components (supraglacial debris, glacial lakes, sea ice, calving fronts, and Antarctic ice‑shelf extent). The benchmark includes multispectral, RGB, and SAR observations from under‑represented regions and tests thirteen geo‑foundation models alongside U‑Net and Vision Transformer baselines. Results show that with frozen encoders U‑Net slightly outperforms TerraMind, but the difference is not statistically significant; fine‑tuning with learning‑rate optimization can dramatically improve performance for some models, while in few‑shot scenarios several foundation models retain over 90 % of their full‑label accuracy.
By Saurabh Kaushik, Lalit Maurya, Beth Tellman, Swalpa Kumar Roy, Valerio Marsocci, Gustau Camps-Valls, Jocelyn Chanussot
The paper "No One Knows the State of the Art in Geospatial Foundation Models" critiques the current lack of standardization in geospatial foundation model (GFM) research, highlighting inconsistencies in evaluation, training, and model release practices across 152 papers. It reports significant discrepancies—46 cross-paper disagreements of at least 10 points for the same model and benchmark, 94 out of 126 papers using unique pretraining configurations, and 39% of papers releasing no model weights. The authors propose six concrete expectations, including named-license weight release, shared core evaluations, and a unified evaluation harness, to address these coordination failures and foster a clearer, comparable understanding of GFM progress.
By Isaac Corley, Nils Lehmann, Caleb Robinson, Gabriel Tseng, Anthony Fuller, Hamed Alemohammad, Evan Shelhamer, Jennifer Marcus, Hannah Kerner
arXiv:2607. 12177v1 Announce Type: new Abstract: The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models.
By Shelley Cazares
arXiv:2608. 00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.
By Fengxiang Wang, Qiuyang Yu, Yueying Li, Mingshuo Chen, Chengchi Fei, Kaiyi Xu, Lixin Gu, Wangxu Wei, Junchao Gong, Lipeng Ma, Jiong Wang, Fenghua Ling, Wenlong Zhang, Xue Yang, Wenjing Yang, Ben Fei, Long Lan