arXiv AI

Emerging Flexible Designs for Geospatial Multimodal Foundation Models

arXiv:2606. 12595v1 Announce Type: cross Abstract: Foundation models are rapidly transforming Earth observation by enabling scalable pretraining across diverse unlabeled geospatial modalities.

arXiv AI
Jun 19

TerraMind: Large-Scale Generative Multimodality for Earth Observation

arXiv:2504. 11171v5 Announce Type: replace-cross Abstract: We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO).

By Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Long\'ep\'e
arXiv AI
Aug 10

SLED: Scalable Location Encoding via Distillation

arXiv:2608. 06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so.

By Kevin Lane, Zhongying Wang, Esther Rolf, Morteza Karimzadeh
arXiv Computer Vision
Sep 7

MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation

MEOX is a compact multimodal masked autoencoder designed for Earth Observation that uses a 2.939 million‑parameter encoder and 3.115 million total parameters. It incorporates sensor‑specific adapters, explicit validity signals, and a shared sparse‑expert block to maintain modality‑dependent processing before a learned patch‑wise fusion, followed by fourteen encoder blocks that process a single spatial sequence with four metadata tokens. Pretrained on 1.228 million MMEarth64 samples, MEOX achieves strong performance on GEO‑Bench tasks, surpassing prior CSMoE results, and demonstrates effective sensor‑flexible representation learning with a modest parameter budget.

By Mohanad Albughdadi
arXiv AI
Sep 1

A Composition-Aware Pretraining Framework for Geospatial Foundation Models

The paper introduces a composition‑aware pretraining framework for geospatial foundation models that explicitly encodes fractional land‑cover mixtures as histogram targets for each satellite image cell. By using Earth Mover’s Distance to distill these composition targets into a 36.8 M‑parameter backbone, the authors demonstrate significant improvements on region‑level tasks such as zero‑shot image retrieval and scene classification, while maintaining competitive performance on fine‑grained tasks like segmentation and object detection. The method outperforms larger models (SatMAE and Prithvi‑EO‑2.0) and achieves a 55.6 % relative boost on the ForestNet‑12 dataset, evidencing the benefit of explicit composition modeling.

By Aryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath, Shrutilipi Bhattacharjee
Hugging Face Trending Papers
Aug 11

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.