arXiv AI

Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

arXiv:2607. 24336v1 Announce Type: new Abstract: City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed across trips and regions.

Hugging Face Trending Papers
Sep 2

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

UTP‑Bench is a new benchmark that evaluates large language models on uncertainty‑aware travel planning, incorporating real‑world data from 504 Indian cities and empirical delay and crowd patterns. It introduces three metrics—Buffer Adequacy Score, Crowd‑Aware Timing Score, and Transport Delay Absorption Score—to measure how well generated itineraries remain robust under stochastic conditions. Experiments show that current LLMs lag behind human planners in buffering, delay‑aware scheduling, and crowd sensitivity.

arXiv Machine Learning
Jun 17

Public transit gains and spatially uneven travel demand changes after NYC congestion pricing

arXiv:2606. 17530v1 Announce Type: cross Abstract: New York City implemented the nation's first cordon-based congestion pricing program in January 2025, providing an opportunity to evaluate how system-wide urban mobility responds to large-scale pricing interventions.

By Donghang Li, Dingyi Zhuang, Yunlin Li, Chenan Shen, Nina Cao, Yunhan Zheng, Shenhao Wang, Jinhua Zhao
arXiv Machine Learning
Jun 2

Scalable Ride-Sourcing Vehicle Rebalancing with Service Accessibility Guarantee: A Constrained Mean-Field Reinforcement Learning Approach

arXiv:2503. 24183v3 Announce Type: replace Abstract: The expansion of ride-sourcing services such as Uber and Lyft has reshaped urban transportation by offering flexible, on-demand mobility via mobile applications.

By Matej Jusup, Kenan Zhang, Zhiyuan Hu, Barna P\'asztor, Andreas Krause, Francesco Corman
arXiv AI
Sep 3

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates large language models on their ability to generate robust itineraries under real-world stochastic conditions. The dataset covers 504 Indian cities, incorporating attractions, restaurants, accommodations, and multi-modal transportation networks, and includes empirical delay distributions and crowd-density patterns to simulate realistic disruptions. Three new metrics—Buffer Adequacy Score, Crowd-Aware Timing Score, and Transport Delay Absorption Score—measure how well generated plans maintain robustness against transit delays and crowd variability, revealing significant gaps between state-of-the-art LLMs and human-authored itineraries.

By Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana
arXiv AI
Jun 2

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.

By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
arXiv AI
Sep 11

HiRAD: A Flexible Large-Scale AGV Routing System

HiRAD is a hierarchical reinforcement learning framework designed for continuous-space routing of large-scale AGV fleets, offering real-time guarantees. It introduces a step-level spatiotemporal representation, separates heading selection from velocity control to shrink the action space, and employs an asynchronous event-driven decision pipeline that reduces inference complexity from O(n²) to O(n) and cuts per-step latency by up to 71%. Experiments on random graphs and two warehouse maps show that HiRAD decreases makespan by 45% to 63% and shortens overall runtime.

By Yunjie Huang, Ruizhong Wu, Mengxuan Zhang, Frodo Kin Sun Chan, Yan Nei Law, Lei Li