arXiv Machine Learning

IBAD: Interpretable Behavioral Anomaly Detection on Human Mobility Data

arXiv:2606. 16023v1 Announce Type: new Abstract: Human mobility appears highly diverse, yet much of a person's daily mobility can be explained by a small set of recurring behavioral templates, such as commuting, school-centered activities, caregiving, nightlife, or errand patterns.

arXiv Machine Learning
Sep 22

Modelling daily activity patterns from mobile phone location data via deep representation learning

The paper introduces the Activity Chain Encoder (ACE), a self‑supervised deep learning model that transforms passively collected mobile phone location data into daily activity representations. ACE integrates pre‑trained urban embeddings, visit timing, and duration, using a Transformer to capture the sequential structure of stays, and is trained via masked activity modelling and contrastive learning without explicit activity labels. The resulting user‑level profiles are clustered and interpreted with temporal‑functional patterns and Census demographics, revealing six distinct weekday activity‑pattern groups in London that differ in daily rhythms, urban contexts, and demographic characteristics.

By Xinglei Wang, Junyuan Liu, Guangsheng Dong, Zichao Zeng, Stephen Law, James Haworth, Tao Cheng
arXiv AI
Jun 9

Mobility-Embedded POIs: Learning What A Place Is and How It Is Used from Human Movement

arXiv:2601. 21149v3 Announce Type: replace-cross Abstract: Recent progress in geospatial foundation models highlights the importance of learning general-purpose representations for real-world locations, particularly points-of-interest (POIs) where human activity concentrates.

By Maria Despoina Siampou, Shushman Choudhury, Shang-Ling Hsu, Neha Arora, Cyrus Shahabi
arXiv AI
Jul 10

MobiDiff: Semantic-Aware Multi-Channel Discrete Diffusion for Human Mobility Data Generation

arXiv:2607. 08357v1 Announce Type: new Abstract: Human mobility data are essential for transportation optimization, urban planning, and resource allocation, yet real-world mobility data are costly to collect and difficult to share due to privacy concerns.

By Rongchao Xu, Lin Jiang, Dahai Yu, Ximiao Li, Taichi Liu, Desheng Zhang, Yuan Tian, Guang Wang
arXiv AI
Aug 19

MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale

MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.

By Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
arXiv Machine Learning
Sep 2

Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

The study investigates whether large language models (LLMs) can predict neighborhood-level human mobility without training data. Using anonymized Cuebiq data across four U.S. metropolitan areas, the authors compare zero‑shot LLM predictions to supervised baselines for various mobility outcomes and assess structural alignment with empirical trends. Results show supervised models outperform LLMs (average accuracy 0.580 vs. 0.435), with LLMs relying on coarse, stable priors that may exhibit biased treatment of protected-group predictors.

By Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya, Naman Awasthi, Kazi Tasnim Zinat, Vanessa Frias-Martinez
arXiv Machine Learning
Sep 22

LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling

LE4Mob is a new location embedding framework that learns inductive, distance‑aware representations from geographic context, enabling it to encode unseen locations and preserve spatial relationships. It builds on contrastive language‑location pre‑training and adds a regularisation objective that encourages the embedding space to reflect geographic distance. Experiments on next‑location prediction and commuter flow generation across multiple datasets show that LE4Mob outperforms strong baselines, especially in inductive settings and when downstream models use direct interactions between location embeddings.

By Xinglei Wang, Stephen Law, Zichao Zeng, Junyuan Liu, Guangsheng Dong, Tao Cheng
arXiv Machine Learning
Sep 7

BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation

BER-PEF is a Bayes‑error‑rate‑based framework that transforms BER estimation into human mobility predictability estimation, enabling comparison of different estimators even when ground truth predictability is not observable. It maps various data types—symbolic sequences, numeric trajectories, contextual features, and learned representations—into a shared feature–label space and evaluates estimator outputs along controlled perturbation curves against a common predictability reference interval. Experiments on datasets such as Foursquare NYC/TKY, GeoLife, and T‑Drive show that several BER‑based estimators outperform existing methods on symbolic sequences and numeric trajectories, and that aggregating evidence across multiple perturbation levels yields a more reliable basis for selecting estimators.

By En Xu, Jingtao Ding, Zhiwen Yu, Yong Li
arXiv Machine Learning
Sep 3

TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis

TrajMind is a framework for diagnosing collective anomalies in urban trajectory data. It separates continuous screening from on-demand diagnosis, using a fast text-only path for alerts and a slow vision‑language path that chains role‑specialized LoRA adapters for detailed, evidence‑backed what‑who‑where‑when records. Experiments show the slow path outperforms baselines by over 15 percentage points in typing and 13 in localization, while the fast path cuts latency by 41% and retains high accuracy.

By Jiahao Wu, Zhenqun Yang, Chen Jason Zhang, Qing Li