arXiv AI By Weiwei Sun, Xuhui Zhou, Weihua Du, Xingyao Wang, Sean Welleck, Graham Neubig, Maarten Sap, Yiming Yang

Training Proactive and Personalized LLM Agents

Read the original on arXiv AI →

The paper proposes shifting AI agent training from isolated task completion to collaborative interaction, defining three key dimensions—Productivity, Proactivity, and Personalization (PPP). It introduces UserVille, an environment with LLM-based user simulators and user-centric feedback, and a multi-objective reinforcement learning framework that optimizes PPP using rewards from task outcomes, question effort, and preference adherence. Experiments on SWE-Bench and BrowseComp-Plus show PPP-trained agents outperform strong LLM baselines, ask more targeted questions, and generalize to unseen preferences and tasks, with a user study underscoring the value of user-centric feedback for effective, supervised collaboration.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

arXiv:2602. 11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications.

By Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao
arXiv AI
Sep 18

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.

By Wenjie Liao, Liangjie Zhao, Zehong Cao