arXiv AI By Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang, Dongyeop Kang

Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

Read the original on arXiv AI →

The paper introduces a framework for designing proactive large‑language‑model agents, centered on three joint principles—Task Capability, Temporal Allocation, and Trust—alongside a five‑dimensional design space. It proposes PROACTIVITY‑GYM, a simulation testbed for evaluating proactive assistance across multi‑day scenarios, and presents empirical findings that highlight performance gaps and the importance of aligning interventions with user trust. Human studies show that misaligned interventions can sharply reduce trust, even when outcomes are correct, underscoring the need for careful joint optimization of the 3T principles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

The paper presents a unified framework for proactive service agents, defining proactivity as an agent’s ability to infer service opportunities from incomplete signals and decide whether to remain silent, ask, assist, or act. It models this as a partially observable sequential decision process constrained by authorization and risk, integrating timing, content, and delivery into a single structured action. The authors categorize existing methods along a decision pipeline—state and need estimation, intervention gating, action construction, and feedback adaptation—and propose standardized metrics for evaluating triggering, timing, calibration, user burden, safety, and policy value across diverse interaction modalities.

By Yan Tang, Tingyu Cao, Yuanbo Tang, Huaze Tang, Keer Hu
arXiv Machine Learning
Sep 14

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.

By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
Jun 4

Unlocking Proactivity in Task-Oriented Dialogue

arXiv:2605. 22240v2 Announce Type: replace Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the conversation toward acceptance within a bounded number of turns.

By Azure Zhang, Ning Gao, Yuqin Dai, Ruiyuan Wu, Jinpeng Wang, Rena Wei Gao, Bingdong Tan, Shuzheng Gao, Zongjie Li, Chaozheng Wang