arXiv AI

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

arXiv AI
Sep 24

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

The paper introduces the concept of an agent’s "taste"—its ability to make effective long‑horizon decisions—and presents Taste‑Bench, a new benchmark that automatically generates decision‑fork questions from agent trajectories. Taste‑Bench evaluates models on choosing the best path without seeing future outcomes, revealing that top models answer only about 60% of questions correctly and that later‑appearing evidence makes forks harder. The authors also demonstrate that training a student model to mimic a teacher’s judgment improves decision quality and overall success on held‑out software engineering tasks.

By Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia
arXiv Computation and Language
Sep 16

Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents

The paper introduces ATRBench, a benchmark that measures the proactivity gap in long‑lived LLM agents by evaluating their ability to ask for user preferences that are not needed immediately but may be useful in future sessions. It defines the Ask‑to‑Remember (ATR) task, where agents must decide whether to request a reusable preference now, and shows that current state‑of‑the‑art agents perform significantly below an oracle. The study identifies preference acquisition as the main bottleneck and provides a diagnostic framework for improving agent proactivity.

By Bin Wu, Guanyun Zou, Bingbing Wang, Huan Zhao, Chuan Shi
arXiv AI
Aug 11

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

arXiv:2608. 07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain.

By Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani