The paper introduces a multimodal foundation model for lunar remote sensing, trained from scratch on SomBench—a dataset of nearly two million co‑registered tile bundles across 11 modalities at 1 m and 100 m resolutions. The model extends the TerraMind masked‑token architecture with lunar‑specific features such as explicit acquisition geometry and joint training of two spatial scales, and employs FlexiViT patch embeddings for adaptable patch sizes. Evaluation on crater detection, irregular mare patch segmentation, and polar ice prospectivity regression shows that the pretrained model matches or surpasses ImageNet‑pretrained baselines, with notable label efficiency and effective adaptation via LoRA.
By Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal, Rachel A. Slank, Geoffrey Dawson, Hiyam Debary, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Nikolaos Dionelis, Ankur Kumar, Campbell D. Watson, Manil Maskey, Rebekah I. Dawson-Rigas, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
arXiv:2606. 12595v1 Announce Type: cross Abstract: Foundation models are rapidly transforming Earth observation by enabling scalable pretraining across diverse unlabeled geospatial modalities.
By Philipe Dias, Waqwoya Abebe, Abhishek Potnis, Aristeidis Tsaris, Dan Lu, Xiao Wang, Dalton Lunga
arXiv:2608. 06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so.
By Kevin Lane, Zhongying Wang, Esther Rolf, Morteza Karimzadeh
Genesis is a generative engine that produces fully consistent multi‑scale satellite image pyramids by combining a vertical super‑resolution model with a horizontal mask‑based outpainting model. It addresses the lack of existing methods that can synthesize a complete quadtree from sparse seed tiles at arbitrary zoom levels and positions, ensuring coherence across both scale and space. The authors also release dense500, a comprehensive multi‑scale dataset and evaluation suite, to benchmark this new task.
By Subash Khanal, Yangzhi Cui, Daniel Cher, Eric Xing, Brian Wei, Srikumar Sastry, Nathan Jacobs
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv:2607. 18504v1 Announce Type: cross Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact?
By Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-B{\o}rre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci
Genesis is a generative engine designed to synthesize complete, globally consistent quadtree pyramids for satellite imagery. It tackles the new multi‑scale tile completion task by combining a vertical super‑resolution model with a horizontal mask‑based outpainting model, enabling seamless generation across arbitrary zoom levels and positions. The authors also release dense500, a fully observed multi‑scale dataset, and a suite of pyramid‑level metrics to benchmark performance.
arXiv:2608.29609v1 Announce Type: new
Abstract: Semantic segmentation is a crucial task for understanding Mars, the most Earth-like planet in our solar system. However, it is challenging because the...
By Ming-Han Lee, Chi-Yeh Chen
SIMPLER is a pre‑fine‑tuning method that reduces inference and deployment costs for Earth Observation foundation models by pruning redundant layers. It uses layer‑wise representation similarity on unlabeled task data to identify and remove up to 79% of parameters without requiring gradients, magnitude heuristics, or hyperparameter tuning. Experiments on Prithvi‑EO‑2, TerraMind, and ImageNet‑pretrained ViT‑MAE show that SIMPLER retains 94% of baseline performance while achieving 2.1× faster training and 2.6× faster inference.
By V\'ictor Barreiro, Johannes Jakubik, Francisco Arg\"uello, Dora B. Heras
The paper introduces a composition‑aware pretraining framework for geospatial foundation models that explicitly encodes fractional land‑cover mixtures as histogram targets for each satellite image cell. By using Earth Mover’s Distance to distill these composition targets into a 36.8 M‑parameter backbone, the authors demonstrate significant improvements on region‑level tasks such as zero‑shot image retrieval and scene classification, while maintaining competitive performance on fine‑grained tasks like segmentation and object detection. The method outperforms larger models (SatMAE and Prithvi‑EO‑2.0) and achieves a 55.6 % relative boost on the ForestNet‑12 dataset, evidencing the benefit of explicit composition modeling.
By Aryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath, Shrutilipi Bhattacharjee
arXiv:2606. 10819v1 Announce Type: cross Abstract: RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery.
By Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li
arXiv:2608. 19766v1 Announce Type: cross Abstract: Self-supervised pretraining on remote sensing imagery typically treats all samples as equally informative, despite large variability in geographic and visual structure.
By Daniele Rege Cambrin, Francesco Rossi, Mattia Varile