arXiv Computation and Language

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

arXiv AI
Aug 28

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 instances from a major Chinese travel platform, each with nearly 40 recorded behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leverages these trajectories, external retrieval tools, and internal memory, outperforming GPT‑4.1 and other baselines on the dataset.

By Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Yunhong Wang
arXiv AI
Jun 2

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.

By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
Hugging Face Trending Papers
Aug 27

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred directly from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 Chinese travel‑planning instances, each with an average of 39.8 past behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leveraging behavior trajectories, external retrieval tools, and internal memory, outperforms strong baselines such as GPT‑4.1 on this challenging dataset.

arXiv AI
Sep 18

TripScore: Aligning LLMs for Real-World Travel Planning via Expert-Calibrated Reward

TripScore is a benchmark and evaluation framework for large language models (LLMs) in travel planning, built from real user logs and calibrated with 1,468 pairwise judgments from 203 travel experts. It uses a hierarchical feasibility gate for format and commonsense checks, and a unified point-wise reward that combines soft quality and preference fulfillment. Experiments show that reinforcement learning fine‑tuning, such as GRPO, consistently outperforms other methods when evaluated with TripScore.

By Yincen Qu, Huan Xiao, Feng Li, Gregory Li, Hui Zhou, Xiangying Dai, Xiaoru Dai, Xuan Huang
arXiv AI
Jul 23

ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking

arXiv:2601. 06487v3 Announce Type: replace-cross Abstract: Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles on open-ended agent tasks with vast solution spaces (e.

By Qiang Zhang, Boli Chen, Fanrui Zhang, Ruixue Ding, Shihang Wang, Qiuchen Wang, Yinfeng Huang, Haonan Zhang, Rongxiang Zhu, Pengyong Wang, Ailin Ren, Xin Li, Pengjun Xie, Jiawei Liu, Ning Guo, Jingren Zhou, Zheng-Jun Zha
arXiv AI
Sep 3

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates large language models on their ability to generate robust itineraries under real-world stochastic conditions. The dataset covers 504 Indian cities, incorporating attractions, restaurants, accommodations, and multi-modal transportation networks, and includes empirical delay distributions and crowd-density patterns to simulate realistic disruptions. Three new metrics—Buffer Adequacy Score, Crowd-Aware Timing Score, and Transport Delay Absorption Score—measure how well generated plans maintain robustness against transit delays and crowd variability, revealing significant gaps between state-of-the-art LLMs and human-authored itineraries.

By Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana
arXiv AI
Jun 11

MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios

arXiv:2602. 22638v2 Announce Type: replace Abstract: Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through natural language interaction and tool-mediated decision making.

By Zhiheng Song, Jingshuai Zhang, Chuan Qin, Chao Wang, Chao Chen, Longfei Xu, Kaikui Liu, Xiangxiang Chu, Hengshu Zhu
Hugging Face Trending Papers
Sep 2

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

UTP‑Bench is a new benchmark that evaluates large language models on uncertainty‑aware travel planning, incorporating real‑world data from 504 Indian cities and empirical delay and crowd patterns. It introduces three metrics—Buffer Adequacy Score, Crowd‑Aware Timing Score, and Transport Delay Absorption Score—to measure how well generated itineraries remain robust under stochastic conditions. Experiments show that current LLMs lag behind human planners in buffering, delay‑aware scheduling, and crowd sensitivity.

arXiv AI
Sep 17

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

RideWay is a new benchmark that evaluates ride‑hailing language agents not just on task completion but on interaction efficiency. It introduces the Efficiency Utility metric, which penalizes agents for excessive tool calls and user‑facing turns relative to a task‑specific reference effort, with human preferences used to calibrate the penalties. Across 58 tasks and 24 models, the metric shows that extra dialogue is penalized more heavily than extra tool use, and it achieves high accuracy in distinguishing trajectories that differ in turns but struggles when differences are only in tool calls.

By Qingnuan Han, Boli Fang, Mingzhi Hou, Claire Liu