Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 instances from a major Chinese travel platform, each with nearly 40 recorded behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leverages these trajectories, external retrieval tools, and internal memory, outperforming GPT‑4.1 and other baselines on the dataset.
By Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Yunhong Wang
Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred directly from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 Chinese travel‑planning instances, each with an average of 39.8 past behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leveraging behavior trajectories, external retrieval tools, and internal memory, outperforms strong baselines such as GPT‑4.1 on this challenging dataset.
arXiv:2606. 09115v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers a path to policy improvement from logged data alone, using historical returns or other measurable outcomes as world feedback.
By Lena Krieger, Xuan Zhao, Zhuo Cao, Qin Wang, Hanno Scharr, Ira Assent
RideWay is a new benchmark that evaluates ride‑hailing language agents not just on task completion but on interaction efficiency. It introduces the Efficiency Utility metric, which penalizes agents for excessive tool calls and user‑facing turns relative to a task‑specific reference effort, with human preferences used to calibrate the penalties. Across 58 tasks and 24 models, the metric shows that extra dialogue is penalized more heavily than extra tool use, and it achieves high accuracy in distinguishing trajectories that differ in turns but struggles when differences are only in tool calls.
By Qingnuan Han, Boli Fang, Mingzhi Hou, Claire Liu
AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant search...
arXiv:2607. 06624v1 Announce Type: new Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents.
By Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin, Sergey Nikolenko