arXiv:2509.09710v3 Announce Type: replace-cross
Abstract: This study introduces a Large Language Model (LLM) scheme for generating key attributes of travel diaries in agent-based transportation model...
By Sepehr Golrokh Amin, Devin Rhoads, Fatemeh Fakhrmoosavi, Nicholas E. Lownes, John N. Ivan
UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates large language models on their ability to generate robust itineraries under real-world stochastic conditions. The dataset covers 504 Indian cities, incorporating attractions, restaurants, accommodations, and multi-modal transportation networks, and includes empirical delay distributions and crowd-density patterns to simulate realistic disruptions. Three new metrics—Buffer Adequacy Score, Crowd-Aware Timing Score, and Transport Delay Absorption Score—measure how well generated plans maintain robustness against transit delays and crowd variability, revealing significant gaps between state-of-the-art LLMs and human-authored itineraries.
By Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana
arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.
By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
UTP‑Bench is a new benchmark that evaluates large language models on uncertainty‑aware travel planning, incorporating real‑world data from 504 Indian cities and empirical delay and crowd patterns. It introduces three metrics—Buffer Adequacy Score, Crowd‑Aware Timing Score, and Transport Delay Absorption Score—to measure how well generated itineraries remain robust under stochastic conditions. Experiments show that current LLMs lag behind human planners in buffering, delay‑aware scheduling, and crowd sensitivity.
arXiv:2607. 15552v1 Announce Type: cross Abstract: Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability.
By Himel Dev, Tanmoy Sen, Madhusudan Basak, Bashima Islam
arXiv:2606. 13835v1 Announce Type: cross Abstract: LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobility patterns or merely generate plausible mobility narratives.
By Gustavo H. Santos, Aline Carneiro Viana, Thiago H. Silva