arXiv:2509.09710v3 Announce Type: replace-cross
Abstract: This study introduces a Large Language Model (LLM) scheme for generating key attributes of travel diaries in agent-based transportation model...
By Sepehr Golrokh Amin, Devin Rhoads, Fatemeh Fakhrmoosavi, Nicholas E. Lownes, John N. Ivan
UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates large language models on their ability to generate robust itineraries under real-world stochastic conditions. The dataset covers 504 Indian cities, incorporating attractions, restaurants, accommodations, and multi-modal transportation networks, and includes empirical delay distributions and crowd-density patterns to simulate realistic disruptions. Three new metrics—Buffer Adequacy Score, Crowd-Aware Timing Score, and Transport Delay Absorption Score—measure how well generated plans maintain robustness against transit delays and crowd variability, revealing significant gaps between state-of-the-art LLMs and human-authored itineraries.
By Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana
arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.
By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
UTP‑Bench is a new benchmark that evaluates large language models on uncertainty‑aware travel planning, incorporating real‑world data from 504 Indian cities and empirical delay and crowd patterns. It introduces three metrics—Buffer Adequacy Score, Crowd‑Aware Timing Score, and Transport Delay Absorption Score—to measure how well generated itineraries remain robust under stochastic conditions. Experiments show that current LLMs lag behind human planners in buffering, delay‑aware scheduling, and crowd sensitivity.
arXiv:2607. 15552v1 Announce Type: cross Abstract: Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability.
By Himel Dev, Tanmoy Sen, Madhusudan Basak, Bashima Islam
arXiv:2606. 13835v1 Announce Type: cross Abstract: LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobility patterns or merely generate plausible mobility narratives.
By Gustavo H. Santos, Aline Carneiro Viana, Thiago H. Silva
The paper demonstrates that fine‑tuning large language models (LLMs) on local tourist trajectory data can predict visitor movements under varying conditions. Using 566 trajectories from Wakayama Castle Park, Japan, the authors fine‑tuned Llama‑3.1‑8B, achieving 49.1% accuracy for next point‑of‑interest predictions and maintaining strong performance even on undersampled scenarios such as rainy days. This shows that LLMs can serve as high‑fidelity, context‑aware behavior models for tourist prediction and enable counterfactual analysis of mobility interventions.
By Tatsuya Amano, Hirozumi Yamaguchi
arXiv:2608. 15867v1 Announce Type: cross Abstract: Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging.
By Farbod Abbasi, Zachary Patterson, Bilal Farooq
arXiv:2606. 12657v1 Announce Type: new Abstract: Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating realistic synthetic trajectory generation.
By Siyu Li, Toan Tran, Lingyi Zhao, Khurram Shafique, Li Xiong
The paper evaluates how Retrieval-Augmented Generation (RAG) can improve Large Language Model (LLM) predictions of travel mode choice. Four retrieval strategies—basic RAG, balanced retrieval, cross‑encoder re‑ranking, and a combination of balanced retrieval with cross‑encoder—are tested on three LLMs (GPT‑4o, o4‑mini, o3) using 2023 Puget Sound travel survey data. Results show that RAG boosts accuracy across models, with GPT‑4o plus balanced retrieval and cross‑encoder achieving 80.8% accuracy, surpassing traditional statistical and machine learning baselines and demonstrating strong zero‑shot transfer.
By Yiming Xu, Junfeng Jiao
Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 instances from a major Chinese travel platform, each with nearly 40 recorded behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leverages these trajectories, external retrieval tools, and internal memory, outperforming GPT‑4.1 and other baselines on the dataset.
By Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Yunhong Wang
CityReal is a modular framework that uses large language model agents to simulate human-aligned urban behavior. It models agents as intention-driven decision makers who pursue coherent mobility and activity plans, learning habits and preferences over time. By training textual adapters to align agent decisions with observed population statistics, CityReal improves realism at both micro and macro levels and can scale to tens of thousands of agents for analyzing crowd density, place popularity, mobility flows, and well‑being under various urban scenarios.
By Nicolas Bougie, Xiaotong Ye, Narimasa Watanabe