CALM is a reproducible hybrid framework that combines an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. It executes a closed traveler‑day loop and evaluates each generative module against an empirical, reproducible baseline, using the 2024 New York City Citywide Mobility Survey data. The framework demonstrates significant improvements in mode‑choice accuracy, quantifies trade‑offs through live‑LLM ablation, and supports controlled stress testing and deterministic replay of downstream simulations.
By Yezhou Cheng
arXiv:2606. 12657v1 Announce Type: new Abstract: Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating realistic synthetic trajectory generation.
By Siyu Li, Toan Tran, Lingyi Zhao, Khurram Shafique, Li Xiong
arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.
By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
arXiv:2608. 20320v1 Announce Type: new Abstract: Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately.
By Narges Ahmadi (McGill University), Yubo Jiao (McGill University), J\^onatas Augusto Manzolli (McGill University), Jiangbo Yu (McGill University), Luis Miranda-Moreno (McGill University)
arXiv:2609.08288v1 Announce Type: new
Abstract: Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly...
By Zijian Shen, Bin Zhou, Jiguang Wang, Ya Zhao, Jintao Ke
The paper evaluates how Retrieval-Augmented Generation (RAG) can improve Large Language Model (LLM) predictions of travel mode choice. Four retrieval strategies—basic RAG, balanced retrieval, cross‑encoder re‑ranking, and a combination of balanced retrieval with cross‑encoder—are tested on three LLMs (GPT‑4o, o4‑mini, o3) using 2023 Puget Sound travel survey data. Results show that RAG boosts accuracy across models, with GPT‑4o plus balanced retrieval and cross‑encoder achieving 80.8% accuracy, surpassing traditional statistical and machine learning baselines and demonstrating strong zero‑shot transfer.
By Yiming Xu, Junfeng Jiao
The paper demonstrates that fine‑tuning large language models (LLMs) on local tourist trajectory data can predict visitor movements under varying conditions. Using 566 trajectories from Wakayama Castle Park, Japan, the authors fine‑tuned Llama‑3.1‑8B, achieving 49.1% accuracy for next point‑of‑interest predictions and maintaining strong performance even on undersampled scenarios such as rainy days. This shows that LLMs can serve as high‑fidelity, context‑aware behavior models for tourist prediction and enable counterfactual analysis of mobility interventions.
By Tatsuya Amano, Hirozumi Yamaguchi
UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates large language models on their ability to generate robust itineraries under real-world stochastic conditions. The dataset covers 504 Indian cities, incorporating attractions, restaurants, accommodations, and multi-modal transportation networks, and includes empirical delay distributions and crowd-density patterns to simulate realistic disruptions. Three new metrics—Buffer Adequacy Score, Crowd-Aware Timing Score, and Transport Delay Absorption Score—measure how well generated plans maintain robustness against transit delays and crowd variability, revealing significant gaps between state-of-the-art LLMs and human-authored itineraries.
By Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana
arXiv:2608.30924v1 Announce Type: new
Abstract: Travel itinerary generation requires balancing strict spatio-temporal constraints with human preferences. Existing LLM-based planners mainly rely on st...
By Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick, Abhik Jana, Shreya Ghosh, Manish Gupta
arXiv:2606. 13835v1 Announce Type: cross Abstract: LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobility patterns or merely generate plausible mobility narratives.
By Gustavo H. Santos, Aline Carneiro Viana, Thiago H. Silva
UTP‑Bench is a new benchmark that evaluates large language models on uncertainty‑aware travel planning, incorporating real‑world data from 504 Indian cities and empirical delay and crowd patterns. It introduces three metrics—Buffer Adequacy Score, Crowd‑Aware Timing Score, and Transport Delay Absorption Score—to measure how well generated itineraries remain robust under stochastic conditions. Experiments show that current LLMs lag behind human planners in buffering, delay‑aware scheduling, and crowd sensitivity.
arXiv:2606. 00572v1 Announce Type: new Abstract: Passenger count data from public transit systems reveals urban mobility patterns and is essential for planning, operation, and optimisation.
By Oluwaleke Yusuf, Adil Rasheed, Frank Lindseth