MAGIC introduces a dense‑reward reinforcement learning framework for generating mixed‑granularity agent graphs in large‑language‑model based multi‑agent systems. The method sequentially selects functional roles, instantiates them as either single agents or reusable groups, and connects them to existing units, optimizing the construction policy with intermediate feedback from probe‑based utility and structural signals. Experiments show MAGIC outperforms state‑of‑the‑art baselines on eight benchmarks and achieves strong inference efficiency.
By Kairui Yang, Ziheng Yi, Xunkai Li, Minghao An, Zhanke Liu, Zekai Chen, Rong-Hua Li
The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.
By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv:2605. 31289v2 Announce Type: replace-cross Abstract: Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL).
By Amir Esterhuysen, Anders Jonsson
arXiv:2607. 17924v1 Announce Type: cross Abstract: Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL).
By Zijian Zhao, Sen Li
The paper introduces DRG-MAPPO, a hierarchical multi‑agent reinforcement learning framework for cooperative air combat. It combines graph‑based relational modeling with dynamic role assignment, using a high‑level policy to allocate tactical roles such as leader and supporter, and a low‑level policy to execute maneuver actions. The approach includes a target‑priority auxiliary task and achieves an 87% win rate in experiments, indicating effective coordination and stability.
By Junlin Liu, Chengwei Li, Yang Gao, Hui Chang, Xinchen Zhang, Zhijun Zhao, Hao Zhao
The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that splits trajectory rewards into per‑subtask shares before policy updates, replacing the scalar advantage used in Group Relative Policy Optimization (GRPO). RLDS employs Subtask‑Decomposed Advantage Estimation (SDAE) to compute group‑relative advantages and distribute credit to tokens based on subtask importance, focusing on steps where a reflection marks a subtask as consequential. Experiments on four benchmarks—FrozenLake, HotpotQA, ScienceWorld, and DeepResearch—show that RLDS improves performance on high‑heterogeneity tasks (ScienceWorld and FrozenLake) and is more compute‑efficient than scalar GRPO for long rollouts.
By Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich