arXiv AI

From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost

arXiv AI
Sep 17

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

RideWay is a new benchmark that evaluates ride‑hailing language agents not just on task completion but on interaction efficiency. It introduces the Efficiency Utility metric, which penalizes agents for excessive tool calls and user‑facing turns relative to a task‑specific reference effort, with human preferences used to calibrate the penalties. Across 58 tasks and 24 models, the metric shows that extra dialogue is penalized more heavily than extra tool use, and it achieves high accuracy in distinguishing trajectories that differ in turns but struggles when differences are only in tool calls.

By Qingnuan Han, Boli Fang, Mingzhi Hou, Claire Liu
arXiv AI
Sep 4

Efficient Test-Time Adaptation through Human-AI Interaction

The paper introduces Test-Time Adaptation through Human‑Agent Interaction (TAHI), a method that uses iterative human feedback to adapt AI agents to individual users’ criteria. By integrating cross‑session interaction data into agent context and weights, and building an evolving rubric module, the authors demonstrate that agents can improve task success by 4.5–20.9% after only a few interactions. The evolving rubric also serves as a scalable annotation tool, detecting 16.0–22.3% more failures than language models or humans alone, and personalized agents can even generalize improvements up to 8.8% across users.

By Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan, Qianou Ma, Luxi He, Zhoujun Cheng, Andre He, Seungone Kim, Jiayi Geng, Mingqian Zheng, Weiwei Sun, Zheyuan Zhang, Xinran Zhao, Yike Wang, Abe Hou, Liwei Jiang, Pang Wei Koh, Diyi Yang, Graham Neubig, Daniel Fried
arXiv AI
Sep 2

AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning

The article argues that conversational AI should provide contingent feedback—responses that vary with user behavior and its social consequences—rather than merely seeking user approval and fluency. It highlights how current alignment methods, such as reinforcement learning from human feedback, often produce sycophantic, noncontingent affirmation, which can hinder the development of interpersonal skills, especially in adolescents. The authors propose a framework for evaluating and designing contingent AI, incorporating trajectory-based assessment and social consequence prediction, and call for interdisciplinary research to ensure AI systems positively influence human social learning.

By Scott Compton, Arjun Nagendran
arXiv AI
Sep 25

IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

IDRBench is a benchmark designed to evaluate the interactive capabilities of deep research agents that use large language models. It introduces controlled opportunities for clarification within a common workflow, comparing autonomous and interactive trajectories by measuring task‑specific report alignment and interaction cost. Experiments on 100 tasks with seven LLMs show that interaction consistently improves alignment, though its effectiveness varies depending on the agents’ questions and feedback integration.

By Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung