Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces a framework for designing proactive large‑language‑model agents, centered on three joint principles—Task Capability, Temporal Allocation, and Trust—alongside a five‑dimensional design space. It proposes PROACTIVITY‑GYM, a simulation testbed for evaluating proactive assistance across multi‑day scenarios, and presents empirical findings that highlight performance gaps and the importance of aligning interventions with user trust. Human studies show that misaligned interventions can sharply reduce trust, even when outcomes are correct, underscoring the need for careful joint optimization of the 3T principles.
arXiv:2607. 02975v1 Announce Type: new Abstract: Effective agency in social environments depends on when an agent seeks knowledge, when it acts, and whether its actions are justified by acquired information.
The paper presents a unified framework for proactive service agents, defining proactivity as an agent’s ability to infer service opportunities from incomplete signals and decide whether to remain silent, ask, assist, or act. It models this as a partially observable sequential decision process constrained by authorization and risk, integrating timing, content, and delivery into a single structured action. The authors categorize existing methods along a decision pipeline—state and need estimation, intervention gating, action construction, and feedback adaptation—and propose standardized metrics for evaluating triggering, timing, calibration, user burden, safety, and policy value across diverse interaction modalities.
S3Gym is an interactive benchmark designed to evaluate large language models (LLMs) on their ability to self-improve through self-testing, self-judging, and self-improvement. It separates permissive exploration from strict held-out evaluation across seven text-based games with executable environment verifiers. Experiments show that self-improvement varies by task, with different experience incorporation pathways (direct history, summary memory, or parameter training) yielding mixed results and highlighting the need for agents to transform feedback into executable, transferable policies.
GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.
arXiv:2609.22592v1 Announce Type: cross Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a ve...