The paper introduces a multimodal foundation model for lunar remote sensing, trained from scratch on SomBench—a dataset of nearly two million co‑registered tile bundles across 11 modalities at 1 m and 100 m resolutions. The model extends the TerraMind masked‑token architecture with lunar‑specific features such as explicit acquisition geometry and joint training of two spatial scales, and employs FlexiViT patch embeddings for adaptable patch sizes. Evaluation on crater detection, irregular mare patch segmentation, and polar ice prospectivity regression shows that the pretrained model matches or surpasses ImageNet‑pretrained baselines, with notable label efficiency and effective adaptation via LoRA.
By Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal, Rachel A. Slank, Geoffrey Dawson, Hiyam Debary, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Nikolaos Dionelis, Ankur Kumar, Campbell D. Watson, Manil Maskey, Rebekah I. Dawson-Rigas, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
arXiv:2609.22379v1 Announce Type: new
Abstract: High-resolution orbital imagery offers a rich record of the Martian surface, but sparse geological labels limit supervised representation learning. We...
By Akshay Naik, Marius F. R. Juston, Jay Mahajan
arXiv:2609.13332v1 Announce Type: new
Abstract: Automated landslide segmentation on Mars is one of the important tasks for understanding its surface processes, and all will aid in future space explor...
By Leo Thomas Ramos, Sidike Paheding, Abel A. Reyes-Angulo, Rajaneesh A., Sajinkumar K. S., Angel D. Sappa, Thomas Oommen
arXiv:2607. 03644v1 Announce Type: cross Abstract: Decades of orbital missions have produced multi-modal remote sensing data for the Moon, spanning optical imagery, spectroscopy, thermal emission, radar, gravity, and elemental composition.
By Ayush Prasad, Swarnalee Mazumder
The paper proposes a method to adapt the Depth Anything V2 (DAV2) zero‑shot relative depth model for estimating lunar surface height. By fine‑tuning DAV2 with publicly available stereophotogrammetry‑derived DEM data, the authors achieve a significant performance boost over the unadapted zero‑shot model. This improved estimator can provide more accurate relative height information useful for hazard detection in future ESA lunar landings.
By Patrick Bauer, Marius Schwinning, Melanie Siegel, Andreas Weinmann, Hichem Snoussi
arXiv:2608.29609v1 Announce Type: new
Abstract: Semantic segmentation is a crucial task for understanding Mars, the most Earth-like planet in our solar system. However, it is challenging because the...
By Ming-Han Lee, Chi-Yeh Chen
arXiv:2606. 22649v2 Announce Type: replace-cross Abstract: Foundation models provide highly descriptive representations for medical images, yet their reliability degrades under distribution shifts arising from changes in patients, devices, or acquisition conditions.
By Francesco Di Salvo, Sebastian Doerrich, Christian Ledig
arXiv:2609.13277v1 Announce Type: cross
Abstract: Lunar orbital missions, such as Lunar Reconnaissance Orbiter, Kaguya/SELENE, Gravity Recovery and Interior Laboratory, and Lunar Prospector, among ot...
By Himanshu Patil, Gabby Nyirjesy, Rachel A. Slank, Vishal Gaur, Daniela Szwarcman, Paolo Fraccaro, Nikolaos Dionelis, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Hiyam Debary, Ankur Kumar, Rohit Lal, Geoffrey Dawson, Campbell Watson, Rebekah I. Dawson-Rigas, Manil Maskey, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
MANTLE is a multi‑task adaptive network designed for planetary perception, featuring a shared DINOv2 backbone with separate heads for landform classification and boulder segmentation. Trained on HiRISE and MSL imagery, it achieved 92.56% accuracy on seven Martian terrain classes and a 0.753 IoU for boulder segmentation, with strong cross‑sol generalization. The framework follows the Modular Uplink Principle, allowing lightweight task‑specific heads to be trained on Earth and uplinked to the rover without retraining the core model.
By Pranav Durai, Gary Doran
SIMPLER is a pre‑fine‑tuning method that reduces inference and deployment costs for Earth Observation foundation models by pruning redundant layers. It uses layer‑wise representation similarity on unlabeled task data to identify and remove up to 79% of parameters without requiring gradients, magnitude heuristics, or hyperparameter tuning. Experiments on Prithvi‑EO‑2, TerraMind, and ImageNet‑pretrained ViT‑MAE show that SIMPLER retains 94% of baseline performance while achieving 2.1× faster training and 2.6× faster inference.
By V\'ictor Barreiro, Johannes Jakubik, Francisco Arg\"uello, Dora B. Heras
Large 3D foundation models such as MASt3R achieve state-of-the-art stereo reconstruction but are computationally demanding for deployment under strict hardware constraints -- a critical limitation in domains such as planetary exploration, where onboard computing is severely restricted. We study how far such models can be compressed through knowledge distillation, using lunar stereo reconstruction as a challenging and practically relevant case study.
The paper presents a method to replace traditional window-based tracking in geostationary atmospheric motion vector (AMV) stereo matching with deep optical flow, enabling efficient and accurate retrieval of dense three‑dimensional wind fields. A stereo teacher model is distilled into a single‑satellite student model that emulates the teacher’s uncertainty estimates, allowing global wind generation from full‑disk GEO imagery. Validation against radiosondes, operational AMVs, ERA5 reanalysis, and EarthCARE cloud profiles shows that the stereo winds outperform operational AMVs in water‑vapor bands while performing slightly worse in the long‑wave infrared band.
By Thomas J. Vandal, Dong L. Wu, James L. Carr, Derek J. Posselt, Elise Penn, Tristan Ballard, August Posch, Kate Duffy