arXiv Machine Learning

MIND the Gap: A Geographic Implicit Neural Representation with Adjustable Spatial Scale

The paper introduces MIND, a method that distills specialist geospatial model embeddings into a single generalist coordinate embedding with adjustable spatial granularity, using nested supervision across multiple embedding dimensions. MIND’s design allows downstream predictors to use only leading chunks or apply a Chunked Penalty to downweight finer details without retraining the INR. The authors evaluate MIND on CoordBench, a large INR benchmark of 52 datasets and 78 targets, and report that MIND and its Chunked Penalty variant achieve the highest regression and classification scores, especially under regional holdout, establishing a new state‑of‑the‑art for geographic implicit neural representations.

arXiv Machine Learning
2d ago

Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales

Location encoders transform geographic coordinates into high‑dimensional embeddings for machine learning, yet it is unclear how well these embeddings capture interpretable spatial effects. This study benchmarks GeoShapley—a game‑theoretic explainer treating all location features as a single joint player—against eleven TorchSpatial encoders on a synthetic process with known coefficients, across grid, county, and global scales, with and without raw coordinates and under different training regimes. The results show that primary coefficient recovery is consistently high across encoders, while secondary coefficient recovery varies more with scale, especially at the global level, and raw‑coordinate baselines remain competitive throughout.

By Daniel Kiv, Shaowen Wang
arXiv AI
Aug 10

SLED: Scalable Location Encoding via Distillation

arXiv:2608. 06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so.

By Kevin Lane, Zhongying Wang, Esther Rolf, Morteza Karimzadeh
arXiv AI
Aug 19

MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale

MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.

By Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
arXiv AI
Jun 19

TerraMind: Large-Scale Generative Multimodality for Earth Observation

arXiv:2504. 11171v5 Announce Type: replace-cross Abstract: We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO).

By Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Long\'ep\'e
arXiv Computer Vision
Sep 24

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.

By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv AI
Jul 20

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

arXiv:2607. 15768v1 Announce Type: cross Abstract: Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space.

By Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu