arXiv AI By Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Read the original on arXiv AI →

arXiv:2608. 12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 2

Iterative GRPO: Batch-Online Policy Iteration for Multi-Turn RL via Single-Turn RLHF

Iterative GRPO is a batch‑online policy iteration framework that enables multi‑turn reinforcement learning for conversational agents without requiring an interactive user simulator. It alternates between learning a turn‑level Q‑function from logged returns (policy evaluation) and applying single‑turn GRPO against this Q‑function (policy improvement), thereby scoring candidate responses by their expected downstream return. The method is validated on six multi‑turn negotiation environments, demonstrating its practicality for real‑world deployment patterns.

By Daniel R. Jiang, Ankur Samanta, Yukai Yang, Jalaj Bhandari, R\'emi Munos, Tyler Lu
arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang