arXiv AI

Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing

arXiv:2607. 03644v1 Announce Type: cross Abstract: Decades of orbital missions have produced multi-modal remote sensing data for the Moon, spanning optical imagery, spectroscopy, thermal emission, radar, gravity, and elemental composition.

arXiv Machine Learning
Jul 27

LunarFM: A Shared Multimodal Representation of the Moon's Surface

arXiv:2607. 22408v1 Announce Type: new Abstract: The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar surface.

By Marc Girona-Mata, Jakob Gawlikowski, Sumit Goski, Gautier Bardi de Fourtou, Valentin T. Bickel, Ben Moseley, Abigail Calzada-Diaz, Sylvester Kaczmarek, Ra\'ul Ramos-Poll\'an
arXiv AI
Sep 15

Multimodal-Multiresolution Foundation Model for Lunar Remote Sensing

The paper introduces a multimodal foundation model for lunar remote sensing, trained from scratch on SomBench—a dataset of nearly two million co‑registered tile bundles across 11 modalities at 1 m and 100 m resolutions. The model extends the TerraMind masked‑token architecture with lunar‑specific features such as explicit acquisition geometry and joint training of two spatial scales, and employs FlexiViT patch embeddings for adaptable patch sizes. Evaluation on crater detection, irregular mare patch segmentation, and polar ice prospectivity regression shows that the pretrained model matches or surpasses ImageNet‑pretrained baselines, with notable label efficiency and effective adaptation via LoRA.

By Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal, Rachel A. Slank, Geoffrey Dawson, Hiyam Debary, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Nikolaos Dionelis, Ankur Kumar, Campbell D. Watson, Manil Maskey, Rebekah I. Dawson-Rigas, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
arXiv Machine Learning
Sep 15

SomBench: Benchmark Dataset for Advancing Machine Learning in Lunar Science

arXiv:2609.13277v1 Announce Type: cross Abstract: Lunar orbital missions, such as Lunar Reconnaissance Orbiter, Kaguya/SELENE, Gravity Recovery and Interior Laboratory, and Lunar Prospector, among ot...

By Himanshu Patil, Gabby Nyirjesy, Rachel A. Slank, Vishal Gaur, Daniela Szwarcman, Paolo Fraccaro, Nikolaos Dionelis, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Hiyam Debary, Ankur Kumar, Rohit Lal, Geoffrey Dawson, Campbell Watson, Rebekah I. Dawson-Rigas, Manil Maskey, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
arXiv Computer Vision
3d ago

Hyperspectral Image Models: Technical Report

The technical report introduces Hyperspectral Image Models, a modular framework that unifies 55 deep‑learning models across six paradigms for hyperspectral remote sensing. It standardizes tensor conventions, evaluation protocols, and dataset handling, integrating 24 benchmark scenes from various sensors and providing tools to avoid train‑test overlap. Experiments across 1,320 model‑scene combinations show that scene difficulty outweighs architecture, with no single paradigm dominating and small models achieving performance comparable to much larger ones.

By Tanishq Rachamalla, Aryan Das, Srishti Kaushik, Swalpa Kumar Roy
arXiv Computer Vision
Sep 7

MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation

MEOX is a compact multimodal masked autoencoder designed for Earth Observation that uses a 2.939 million‑parameter encoder and 3.115 million total parameters. It incorporates sensor‑specific adapters, explicit validity signals, and a shared sparse‑expert block to maintain modality‑dependent processing before a learned patch‑wise fusion, followed by fourteen encoder blocks that process a single spatial sequence with four metadata tokens. Pretrained on 1.228 million MMEarth64 samples, MEOX achieves strong performance on GEO‑Bench tasks, surpassing prior CSMoE results, and demonstrates effective sensor‑flexible representation learning with a modest parameter budget.

By Mohanad Albughdadi
arXiv Computer Vision
Sep 15

Global-Local Contextual Progressive Expansion Network for Martian Landslide Segmentation in Multimodal Remote Sensing Imagery

arXiv:2609.13332v1 Announce Type: new Abstract: Automated landslide segmentation on Mars is one of the important tasks for understanding its surface processes, and all will aid in future space explor...

By Leo Thomas Ramos, Sidike Paheding, Abel A. Reyes-Angulo, Rajaneesh A., Sajinkumar K. S., Angel D. Sappa, Thomas Oommen
arXiv Machine Learning
Jun 10

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.

By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
arXiv AI
Aug 10

SLED: Scalable Location Encoding via Distillation

arXiv:2608. 06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so.

By Kevin Lane, Zhongying Wang, Esther Rolf, Morteza Karimzadeh