From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
RideWay is a new benchmark that evaluates ride‑hailing language agents not just on task completion but on interaction efficiency. It introduces the Efficiency Utility metric, which penalizes agents for excessive tool calls and user‑facing turns relative to a task‑specific reference effort, with human preferences used to calibrate the penalties. Across 58 tasks and 24 models, the metric shows that extra dialogue is penalized more heavily than extra tool use, and it achieves high accuracy in distinguishing trajectories that differ in turns but struggles when differences are only in tool calls.
AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant search...
arXiv:2606. 09833v1 Announce Type: cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work.
The paper introduces Test-Time Adaptation through Human‑Agent Interaction (TAHI), a method that uses iterative human feedback to adapt AI agents to individual users’ criteria. By integrating cross‑session interaction data into agent context and weights, and building an evolving rubric module, the authors demonstrate that agents can improve task success by 4.5–20.9% after only a few interactions. The evolving rubric also serves as a scalable annotation tool, detecting 16.0–22.3% more failures than language models or humans alone, and personalized agents can even generalize improvements up to 8.8% across users.
arXiv:2608. 06329v1 Announce Type: cross Abstract: Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed.
arXiv:2609.07594v1 Announce Type: cross Abstract: Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world tasks such as cooking and DIY through...