arXiv:2607. 10651v1 Announce Type: new Abstract: In large urban areas, planning multi-day travel itineraries is challenging due to the abundance of Points of Interest (POIs), diverse user preferences, and constraints such as opening hours.
By Rongbo Qi, Yaqi Zhang, Shijun Yan, Xuemeng Liu, Xiangrui Cai, Chunyao Song
arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.
By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
TripScore is a benchmark and evaluation framework for large language models (LLMs) in travel planning, built from real user logs and calibrated with 1,468 pairwise judgments from 203 travel experts. It uses a hierarchical feasibility gate for format and commonsense checks, and a unified point-wise reward that combines soft quality and preference fulfillment. Experiments show that reinforcement learning fine‑tuning, such as GRPO, consistently outperforms other methods when evaluated with TripScore.
By Yincen Qu, Huan Xiao, Feng Li, Gregory Li, Hui Zhou, Xiangying Dai, Xiaoru Dai, Xuan Huang
Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 instances from a major Chinese travel platform, each with nearly 40 recorded behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leverages these trajectories, external retrieval tools, and internal memory, outperforming GPT‑4.1 and other baselines on the dataset.
By Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Yunhong Wang
arXiv:2607. 15562v1 Announce Type: cross Abstract: Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits.
By Himel Dev, Madhusudan Basak, Tanmoy Sen, Paromita Shome, Bashima Islam
Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred directly from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 Chinese travel‑planning instances, each with an average of 39.8 past behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leveraging behavior trajectories, external retrieval tools, and internal memory, outperforms strong baselines such as GPT‑4.1 on this challenging dataset.
arXiv:2509. 21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generation, ensuring an enjoyable user experience.
By Yansong Ning, Rui Liu, Jun Wang, Kai Chen, Wei Li, Jun Fang, Kan Zheng, Naiqiang Tan, Hao Liu
arXiv:2608.30924v1 Announce Type: new
Abstract: Travel itinerary generation requires balancing strict spatio-temporal constraints with human preferences. Existing LLM-based planners mainly rely on st...
By Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick, Abhik Jana, Shreya Ghosh, Manish Gupta
The paper introduces PrefDT, a preference-conditioned Decision Transformer designed for multi‑objective scheduling of UAV mobile edge computing fleets. PrefDT accepts a desired energy‑delay trade‑off as input, enabling a single offline‑trained model to generate any point on the Pareto front during runtime. The authors employ attention pooling with a per‑user bypass to maintain scheduler operation when user reports are lost, and a distillation pipeline to create a preference‑labeled flight corpus, achieving superior trade‑off curves and tight energy budget adherence in simulations.
By Qiao Liao, Zhiyong Feng, Bin Wu, Guodong Fan
CALM is a reproducible hybrid framework that combines an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. It executes a closed traveler‑day loop and evaluates each generative module against an empirical, reproducible baseline, using the 2024 New York City Citywide Mobility Survey data. The framework demonstrates significant improvements in mode‑choice accuracy, quantifies trade‑offs through live‑LLM ablation, and supports controlled stress testing and deterministic replay of downstream simulations.
By Yezhou Cheng
UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates large language models on their ability to generate robust itineraries under real-world stochastic conditions. The dataset covers 504 Indian cities, incorporating attractions, restaurants, accommodations, and multi-modal transportation networks, and includes empirical delay distributions and crowd-density patterns to simulate realistic disruptions. Three new metrics—Buffer Adequacy Score, Crowd-Aware Timing Score, and Transport Delay Absorption Score—measure how well generated plans maintain robustness against transit delays and crowd variability, revealing significant gaps between state-of-the-art LLMs and human-authored itineraries.
By Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana
arXiv:2605.25200v3 Announce Type: replace
Abstract: Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single us...
By Xiang Cheng, Yulan Hu, Lulu Zheng, Xiangwen Zhang, Zheng Pan, Xin Li, Yong Liu