arXiv:2609.00454v1 Announce Type: new
Abstract: Pretrained transformer-based language models achieve strong performance across a wide range of NLP tasks but remain limited in encoding geo-locational...
By Gokul Srinivasagan, Munir Georges
arXiv:2606. 24997v1 Announce Type: new Abstract: Geographic implicit neural representations (INRs) learn to map any coordinate on Earth to a location embedding, implicitly encoding geospatial data into the weights of a neural network.
By Livia Betti, Sebastian Ricke, Ivica Obadic, Adam J. Stewart, Esther Rolf
arXiv:2608. 03826v1 Announce Type: cross Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues.
By Jiapeng Li, Yong Li, Junjie Zhou, Fan Zhang, Yu Liu
arXiv:2601. 21149v3 Announce Type: replace-cross Abstract: Recent progress in geospatial foundation models highlights the importance of learning general-purpose representations for real-world locations, particularly points-of-interest (POIs) where human activity concentrates.
By Maria Despoina Siampou, Shushman Choudhury, Shang-Ling Hsu, Neha Arora, Cyrus Shahabi
arXiv:2602. 15029v3 Announce Type: replace Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe.
By Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri
arXiv:2508. 07195v2 Announce Type: replace-cross Abstract: Recent advances have demonstrated that Large Language Models (LLMs) can be effectively adapted for time series forecasting, revealing strong potential beyond natural language tasks.
By Yanru Sun, Emadeldeen Eldele, Zongxia Xie, Yucheng Wang, Wenzhe Niu, Qinghua Hu, Chee Keong Kwoh, Min Wu
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes.
The paper introduces a multi‑view contrastive learning framework that creates low‑dimensional spatial embeddings by combining satellite imagery and OpenStreetMap data across Europe. These embeddings align with coordinate‑based encodings, allowing any dataset with latitude‑longitude pairs to be enriched with meaningful spatial features without needing the original spatial inputs. In case studies on French real‑estate prices and Belgian flood claim counts, models using the embeddings outperform those using raw coordinates, improving predictive accuracy and territorial risk classification while offering explainable spatial effects.
By Freek Holvoet, Christopher Blier-Wong, Katrien Antonio
arXiv:2606. 28369v1 Announce Type: cross Abstract: Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.
By Yuanyuan Tian, Wenwen Li, Xiao Chen, Michael Brook, Michael Brubaker, Anna Liljedahl, Chitta Baral
ShapeLex introduces a two-stage approach for text-controlled time series generation. It first creates a reusable vocabulary of discrete shape units—such as rises, spikes, and sharp drops—derived from training data, then uses an autoregressive generator to select and arrange these shapes based on textual input while adjusting their position and duration. Finally, a mixture-density scale head models global attributes like level and volatility to produce realistic time series that align closely with real data distributions.
By Subo Wei, Jianqi Gao, Mingyan Fan, Shaorong Xie, Xinzhi Wang, Yongpeng Dong
SAGE is a CLIP-based framework that augments vision‑language time series forecasting by incorporating variable‑specific semantic and statistical information. It processes frequency‑enhanced patches and variable tokens through a CLIP text encoder, while a frozen CLIP vision encoder aligns rendered series with temporal representations via a contrastive objective. The approach achieves state‑of‑the‑art accuracy on eight long‑term benchmarks and M4, with ablations showing complementary gains from multimodal alignment and variable‑level knowledge.
By Haizhao Fan, Xinyi Le
GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.
By Jinnao Li, Tingzhu Chen, Changbo Wang