MAGIC introduces a dense‑reward reinforcement learning framework for generating mixed‑granularity agent graphs in large‑language‑model based multi‑agent systems. The method sequentially selects functional roles, instantiates them as either single agents or reusable groups, and connects them to existing units, optimizing the construction policy with intermediate feedback from probe‑based utility and structural signals. Experiments show MAGIC outperforms state‑of‑the‑art baselines on eight benchmarks and achieves strong inference efficiency.
By Kairui Yang, Ziheng Yi, Xunkai Li, Minghao An, Zhanke Liu, Zekai Chen, Rong-Hua Li
The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.
By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv:2605. 31289v2 Announce Type: replace-cross Abstract: Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL).
By Amir Esterhuysen, Anders Jonsson
arXiv:2607. 17924v1 Announce Type: cross Abstract: Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL).
By Zijian Zhao, Sen Li
The paper introduces DRG-MAPPO, a hierarchical multi‑agent reinforcement learning framework for cooperative air combat. It combines graph‑based relational modeling with dynamic role assignment, using a high‑level policy to allocate tactical roles such as leader and supporter, and a low‑level policy to execute maneuver actions. The approach includes a target‑priority auxiliary task and achieves an 87% win rate in experiments, indicating effective coordination and stability.
By Junlin Liu, Chengwei Li, Yang Gao, Hui Chang, Xinchen Zhang, Zhijun Zhao, Hao Zhao
The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that splits trajectory rewards into per‑subtask shares before policy updates, replacing the scalar advantage used in Group Relative Policy Optimization (GRPO). RLDS employs Subtask‑Decomposed Advantage Estimation (SDAE) to compute group‑relative advantages and distribute credit to tokens based on subtask importance, focusing on steps where a reflection marks a subtask as consequential. Experiments on four benchmarks—FrozenLake, HotpotQA, ScienceWorld, and DeepResearch—show that RLDS improves performance on high‑heterogeneity tasks (ScienceWorld and FrozenLake) and is more compute‑efficient than scalar GRPO for long rollouts.
By Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich
arXiv:2605. 24202v2 Announce Type: replace Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that are poorly understood.
By Yifan Zeng, Yiran Wu, Yaolun Zhang, Wentian Zhao, Kun Wan, Qingyun Wu, Huazheng Wang
arXiv:2609.14968v1 Announce Type: new
Abstract: Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay betwe...
By Tiangang Li, Shi Ying, Xiangbo Tian
arXiv:2606. 25073v1 Announce Type: new Abstract: In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch for each new environment or task.
By Animesh Animesh, Satheesh K Perepu, Kaushik Dey
arXiv:2606. 24958v1 Announce Type: new Abstract: Collective behavior arises when locally interacting units produce coordinated global organization, from synchronization in dynamical systems to task-relevant information flow on graphs.
By Ji Chen, Song Chen, Chengzhang Gong, Li Fan, Chao Xu
arXiv:2511. 13103v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations.
By Vidur Sinha, Muhammed Ustaomeroglu, Guannan Qu
arXiv:2601. 17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observability.
By Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif