arXiv:2502. 00345v2 Announce Type: replace-cross Abstract: The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications.
By Yurui Li, Yuxuan Chen, Xiaoli Yang, Shijian Li, Gang Pan
UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.
By Wenjie Liao, Liangjie Zhao, Zehong Cao
arXiv:2509. 10656v2 Announce Type: replace-cross Abstract: For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning.
By Chirayu Nimonkar, Shlok Shah, Catherine Ji, Benjamin Eysenbach
arXiv:2606. 02337v1 Announce Type: new Abstract: Constrained Multi-agent reinforcement learning (CMARL) faces two intertwined challenges: the joint action space grows exponentially with the number of agents, and additional requirements couple agents in ways that reward structure alone does not capture.
By Santiago Amaya-Corredor, Miguel Calvo-Fullana, Anders Jonsson
arXiv:2511. 02304v2 Announce Type: replace-cross Abstract: We study learning multi-task, multi-agent policies for cooperative, temporal objectives, under centralized training, decentralized execution.
By Beyazit Yalcinkaya, Marcell Vazquez-Chanlatte, Ameesh Shah, Hanna Krasowski, Sanjit A. Seshia
ArenaFlow is a hierarchical credit propagation framework designed to improve reinforcement learning for open-ended agent tasks. It uses tournament-based relative ranking to generate trajectory-level rewards and structured reflective evaluation to identify pivotal success steps, reusable strategy skills, and skill usage attribution. The framework propagates advantages to high-confidence steps and maintains a global skill memory, enabling more targeted optimization and reusable skill priors for future exploration.
By Qiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang, Yinfeng Huang, Yi Zheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha
CollabFlow introduces a recursive self‑improvement framework for multi‑agent collaboration in large language model systems. It trains a Collab‑Director to assemble teams of agents, uses a frozen executor to run them, and retrains the director each round based on outcomes. The system incorporates evidence‑conditioned communication protocols within collaboration graphs and a Collaborative Trajectory Balance objective to maintain diverse high‑performing teams across rounds, achieving superior performance on twelve datasets.
By Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang
arXiv:2608. 14332v1 Announce Type: cross Abstract: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them.
By Anisa Saleem, Duksu Kim
arXiv:2607. 17760v1 Announce Type: cross Abstract: Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations.
By Ziyi Liu, Grace Zhang
arXiv:2601. 21523v2 Announce Type: replace Abstract: To promote cooperation in Multi-Agent Reinforcement Learning, the reward signals of all agents can be aggregated together, forming global rewards that are commonly known as the fully cooperative setting.
By Bang Giang Le, Viet Cuong Ta
MAGIC introduces a dense‑reward reinforcement learning framework for generating mixed‑granularity agent graphs in large‑language‑model based multi‑agent systems. The method sequentially selects functional roles, instantiates them as either single agents or reusable groups, and connects them to existing units, optimizing the construction policy with intermediate feedback from probe‑based utility and structural signals. Experiments show MAGIC outperforms state‑of‑the‑art baselines on eight benchmarks and achieves strong inference efficiency.
By Kairui Yang, Ziheng Yi, Xunkai Li, Minghao An, Zhanke Liu, Zekai Chen, Rong-Hua Li
arXiv:2607. 04972v1 Announce Type: cross Abstract: Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partners, and varying team sizes, yet existing approaches often address these challenges in isolation under the closed-world assumption of fixed teammates.
By Yang Li, Feng Xue, Fan Mo, Yunhao Liu, Jianhong Wang, Ying Wen, Qingrui Zhang, Shaoshuai Mou, Wei Pan