arXiv:2607. 24885v1 Announce Type: cross Abstract: Predicting traffic flow is crucial to optimizing transportation systems and improving urban mobility.
By Jinpeng Chen, Ziyu Yu, Tao Wang, Jun Ma, Hongbo Gao, Senzhang Wang, Zufeng Zhang, Kaimin Wei
LE4Mob is a new location embedding framework that learns inductive, distance‑aware representations from geographic context, enabling it to encode unseen locations and preserve spatial relationships. It builds on contrastive language‑location pre‑training and adds a regularisation objective that encourages the embedding space to reflect geographic distance. Experiments on next‑location prediction and commuter flow generation across multiple datasets show that LE4Mob outperforms strong baselines, especially in inductive settings and when downstream models use direct interactions between location embeddings.
By Xinglei Wang, Stephen Law, Zichao Zeng, Junyuan Liu, Guangsheng Dong, Tao Cheng
MoRAX is a lightweight framework that augments geospatial foundation model embeddings with functional structure derived from human mobility data. By incorporating mobility flows, MoRAX preserves the coverage and consistency of existing geospatial models while adding information about functional connectivity among urban regions, enabling zero‑shot deployment in unseen cities. Experiments across four cities in two countries show that the MoRAX teacher model outperforms baseline geospatial models on eight socioeconomic and environmental prediction tasks, and the student model—without direct mobility input—approaches the teacher’s performance.
By Ya Wen, Jixuan Cai, Yulun Zhou, Alec Kirkley
arXiv:2606. 09086v1 Announce Type: new Abstract: Dynamic origin-destination (OD) flow generation seeks to synthesize realistic mobility dynamics from temporal context alone, without relying on historical OD observations.
By Jie Zhao, Xianqi Dai, Jie Feng, Huandong Wang, Yong Li
MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.
By Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
The paper introduces Mobility Stream-Structure Synergy (MoSS), a method that fuses two complementary views of mobility data—an hourly inflow/outflow Sequence view and a Structure view derived from zigzag persistence diagrams—to capture temporal dynamics and evolving regional connectivity. MoSS employs a synergy module that extracts higher‑order representations from the co‑occurrence of these views, moving beyond additive fusion. Experiments on New York City and Chicago demonstrate that MoSS outperforms existing baselines on three downstream tasks using only mobility data.
By Namwoo Kim, Jeeyun Chang, Kanghoon Lee, Yoonjin Yoon
arXiv:2510. 14819v3 Announce Type: replace-cross Abstract: Trajectory representation learning (TRL) aims to encode raw trajectory data into low-dimensional embeddings for downstream tasks such as travel time estimation, mobility prediction, and trajectory similarity analysis.
By Ji Cao, Yu Wang, Tongya Zheng, Jie Song, Qinghong Guo, Zujie Ren, Canghong Jin, Gang Chen, Mingli Song
arXiv:2606.08918v2 Announce Type: replace
Abstract: Worldwide image geo-localization aims to determine where on Earth a single image was captured. However, visually similar scenes may lie thousands o...
By Junchao Cui, Xuanzi Ma, Wenqi Shi, Nan Wu, Biru Zhu, Xiangyang Luo
GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.
By Jinnao Li, Tingzhu Chen, Changbo Wang
The paper introduces DF-LLM, a Dynamic Fusion Large Language Model designed for traffic flow prediction. It combines a spatiotemporal embedding module, a fusion module that uses graph convolution to capture spatial topology and dynamic dependencies, and an LLM backbone with differentiated parameter adaptation and context aggregation attention. Experiments on four datasets demonstrate that DF-LLM outperforms existing methods in predictive accuracy.
By Xue Qiu, Jianli Xiao
The paper introduces DF-LLM, a Dynamic Fusion Large Language Model designed for traffic flow prediction. It combines a spatiotemporal embedding module, a fusion module that uses graph convolution to capture spatial topology and dynamic dependencies, and an LLM backbone with differentiated parameter adaptation and context aggregation attention. Experiments on four datasets show that DF-LLM outperforms existing methods in predictive accuracy.
arXiv:2606. 08046v1 Announce Type: new Abstract: We present OSMGraphCLIP, a CLIP-style geospatial representation model that learns global location embeddings from freely available OpenStreetMap (OSM) data.
By Dimitrios Michail, Eleni Saka, Ioannis Giannopoulos, Ioannis Papoutsis