The paper introduces Consistent Plan-Act (ConPAct), a method that addresses coordination failures between high-level planners and low-level actors in long-horizon agentic tasks. By prompting both agents to produce structured state assertions and programmatically detecting contradictions, the authors identify a systematic planner-actor state mismatch. ConPAct feeds these detected contradictions back to both agents, fine‑tunes them on consistent interactions, and achieves notable performance gains, such as raising MiniGrid success rates from 38.6% to 54.4% with GPT‑5.6‑sol/terra.
By Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang, Yueqing Sun, Jiayuan Zhang, Qi Gu, Han-Jia Ye
arXiv:2605. 18077v2 Announce Type: replace Abstract: Communication is a key component in multi-agent reinforcement learning (MARL) for mitigating partial observability, yet prior approaches often rely on inefficient information exchange or fail to transmit sufficient state information.
By Sangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul Han
arXiv:2606. 30966v1 Announce Type: new Abstract: Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives.
By Arshia Rafieioskouei, Tzu-Han Hsu, Matthew Lucas, Borzoo Bonakdarpour
arXiv:2607. 00155v1 Announce Type: new Abstract: We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes.
By Yunjin Tong
arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.
By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey
The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.
By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang