arXiv Machine Learning By Anne Josiane Kouam, Hristo Boyadzhiev, Konrad Rieck

Secrets Everywhere: Auditing Memorization in Mobility Prediction Models

Read the original on arXiv Machine Learning →

arXiv:2608. 02052v1 Announce Type: new Abstract: Human mobility prediction models, which forecast the next location in a user's trajectory, are increasingly deployed in urban analytics, navigation, and personalized services.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 2

Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

The study investigates whether large language models (LLMs) can predict neighborhood-level human mobility without training data. Using anonymized Cuebiq data across four U.S. metropolitan areas, the authors compare zero‑shot LLM predictions to supervised baselines for various mobility outcomes and assess structural alignment with empirical trends. Results show supervised models outperform LLMs (average accuracy 0.580 vs. 0.435), with LLMs relying on coarse, stable priors that may exhibit biased treatment of protected-group predictors.

By Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya, Naman Awasthi, Kazi Tasnim Zinat, Vanessa Frias-Martinez
arXiv Computation and Language
Sep 2

Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry

The paper introduces a new privacy vulnerability in diffusion language models (DLMs) called token‑level memorization asymmetry, derived from theoretical analysis of diffusion training dynamics. It proposes Q‑Skew, a quantile‑weighted skewness indicator, to perform membership inference on fine‑tuned DLMs, outperforming existing baselines across multiple datasets and models. Additionally, Q‑Skew can be used to extract personally identifiable information (PII), demonstrating a broader privacy attack surface.

By Shengfang Zhai, Leo Marchyok, Yuling Shi, Huanran Chen, Yinpeng Dong, Jiaheng Zhang, Sanghyun Hong
arXiv AI
Jul 10

MobiDiff: Semantic-Aware Multi-Channel Discrete Diffusion for Human Mobility Data Generation

arXiv:2607. 08357v1 Announce Type: new Abstract: Human mobility data are essential for transportation optimization, urban planning, and resource allocation, yet real-world mobility data are costly to collect and difficult to share due to privacy concerns.

By Rongchao Xu, Lin Jiang, Dahai Yu, Ximiao Li, Taichi Liu, Desheng Zhang, Yuan Tian, Guang Wang
arXiv AI
Aug 7

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

arXiv:2608. 05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.

By Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang