DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.
By Joshua Dimasaka, Christian Gei{\ss}, Emily So
arXiv:2607. 29527v1 Announce Type: cross Abstract: A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth.
By Carlos Rodriguez-Pardo, Massimo Tavoni
arXiv:2609.00661v1 Announce Type: new
Abstract: Satellite foundation models offer a globally available alternative to census data for commuting origin-destination (OD) generation, yet no study has sy...
By Ashiq Shukoor Iqbal, Wilson Wongso, Flora D. Salim
MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.
By Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
arXiv:2607. 15768v1 Announce Type: cross Abstract: Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space.
By Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu
arXiv:2609.38165v1 Announce Type: cross
Abstract: The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However,...
By Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini
The study evaluates how the length of observation windows affects the performance of Tessera embeddings for land‑use/land‑cover mapping. By freezing the encoder and recomputing embeddings from a full year down to a single day, the authors benchmark linear probes and UNet heads on LUCAS, DynamicEarthNet, and PASTIS‑R datasets. Results show that embeddings are highly task‑dependent: for phenology‑driven classes (PASTIS‑R) they outperform from‑scratch models by ~46%, while for temporally stable classes (DynamicEarthNet, LUCAS) they match only with full supervision, yet remain more label‑efficient across all datasets.
By Julia Guerrero-Viu, Alex L\'opez-Cifuentes, Ignacio P\'erez-Villar, Fabio Pacifici
arXiv:2609.39366v1 Announce Type: new
Abstract: Remote sensing object counting estimates how many buildings, vehicles, or ships appear in overhead images. Most supervised counters predict a density m...
By Junjing Zheng, Zhiyi Zhou, Ningrui Yang, Hongying Meng
The paper presents a transfer‑learning approach that adapts a multimodal spatiotemporal vision transformer, originally trained on Demographic and Health Survey data, to estimate socioeconomic conditions in forced‑displacement settings. Using satellite‑derived geospatial covariates, the adapted model explains up to 66% of variation in socioeconomic outcomes in camp‑intersecting grids and 41% in non‑camp areas, achieving mean absolute errors of 4.37 and 5.41 index points respectively. This framework supplements periodic household surveys by providing regularly updated, spatially granular socioeconomic estimates that bridge data gaps between survey rounds.
By Steven Ndung'u, Adel Daoud, Ismael Yacoubou Djima, Hai-Anh H. Dang, Patrick Michael Brock
The study evaluates the use of frozen geospatial foundation embeddings (AlphaEarth) for mapping cultivated versus non‑cultivated land in Maine. Using 192 spatially separated patches and USDA Cropland Data Layer labels, a lightweight classifier achieved 93.7% overall accuracy without fine‑tuning, and a nearest‑class‑centroid rule reached 90.2%. A balanced sample of 60,000 labeled pixels was nearly as effective as the full 8.6 million‑pixel pool, and classifiers trained in one year remained accurate across 2018‑2023. In a blind human validation of 385 points, the AlphaEarth‑plus‑random‑forest map matched 95.3% of the consensus, outperforming the CDL reference (91.7%).
By Mohammad Ammar Mughees, Giovanni Montefoschi, Zhongxin Chen, Maria Antonia Brovelli
The paper introduces MIND, a method that distills specialist geospatial model embeddings into a single generalist coordinate embedding with adjustable spatial granularity, using nested supervision across multiple embedding dimensions. MIND’s design allows downstream predictors to use only leading chunks or apply a Chunked Penalty to downweight finer details without retraining the INR. The authors evaluate MIND on CoordBench, a large INR benchmark of 52 datasets and 78 targets, and report that MIND and its Chunked Penalty variant achieve the highest regression and classification scores, especially under regional holdout, establishing a new state‑of‑the‑art for geographic implicit neural representations.
By Isaac Corley, Arjun Rao, Esther Rolf, Konstantin Klemmer, Evan Shelhamer, Nils Lehmann, Marc Ru{\ss}wurm, Gengchen Mai, Nathan Jacobs, Hannah Kerner
The study adapts a convolutional neural network to classify galaxy morphologies using crowd-sourced annotations from Galaxy Zoo 1. It evaluates how training strategies—such as training all layers versus only the last, incorporating hierarchical labels, varying data volume and annotator agreement, staged transfer learning, and ensembling—affect accuracy and efficiency. Results show that full-network training and high annotator agreement yield over 99% accuracy, while hierarchical approaches and staged learning help when data are limited.
By Luis Enrique Sucar, Carlos del Burgo, Jonathan Serrano-P\'erez