arXiv:2603. 02070v3 Announce Type: replace Abstract: When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI planner according to their preferences and expertise.
By Guilhem Fouilh\'e, Rebecca Eifler, Antonin Poch\'e, Sylvie Thi\'ebaux, Nicholas Asher
arXiv:2606. 29225v1 Announce Type: new Abstract: LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system prompts.
By Seongjae Kang, Taehyung Yu, Sung Ju Hwang
The paper introduces control‑data flow separation to improve prompt optimization in multi‑agent large language model systems. By representing execution protocols as typed, validated program objects and keeping task‑relevant content as unstructured language, the method prevents prompt edits from corrupting critical routing, formatting, or termination signals. Experiments on synthetic reasoning, collaborative review generation, and insurance rating workflows show that this approach maintains 100% protocol validity while consistently enhancing task performance.
By Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti, Utkarsh Soni, Yifan Nie, Eugene Wen, Yuntian Deng
The paper introduces HiPS, a hierarchical strategy co‑evolution framework for memory‑augmented agents that separates memory management into a globally shared foundation and a user‑specific adaptive tier. HiPS uses a Universal Strategy to capture shared principles from cross‑persona trajectories, Persona Delta Distillation to create tailored rules for users deviating from general patterns, and Cross‑Level Rule Flow to dynamically adjust the boundary between global and personal rules. Experiments show that this approach consistently outperforms existing memory‑augmented baselines.
By Yupeng Han, Shuochen Liu, Kai Zhang, Ze Liu, Zhihong Pan, Xianquan Wang
arXiv:2602. 11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications.
By Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao
The paper proposes shifting AI agent training from isolated task completion to collaborative interaction, defining three key dimensions—Productivity, Proactivity, and Personalization (PPP). It introduces UserVille, an environment with LLM-based user simulators and user-centric feedback, and a multi-objective reinforcement learning framework that optimizes PPP using rewards from task outcomes, question effort, and preference adherence. Experiments on SWE-Bench and BrowseComp-Plus show PPP-trained agents outperform strong LLM baselines, ask more targeted questions, and generalize to unseen preferences and tasks, with a user study underscoring the value of user-centric feedback for effective, supervised collaboration.
By Weiwei Sun, Xuhui Zhou, Weihua Du, Xingyao Wang, Sean Welleck, Graham Neubig, Maarten Sap, Yiming Yang