arXiv Machine Learning By Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang

HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

Read the original on arXiv Machine Learning →

HySTAR is a MAPPO-based framework that addresses structural target drift in cooperative multi‑agent reinforcement learning by anchoring an overlapping sparse hypergraph as a stable high‑order value‑decomposition scaffold. It separates adaptive representation learning from a temporally consistent decomposition basis, using a spatiotemporal encoder to capture physical and task‑dependent interactions and combining temporal and structural relevance to compute agent‑specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE show consistent improvements over MAPPO‑style, value‑factorization, and dynamic‑grouping baselines, achieving significant gains in performance and convergence speed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 23

MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning

MAGIC introduces a dense‑reward reinforcement learning framework for generating mixed‑granularity agent graphs in large‑language‑model based multi‑agent systems. The method sequentially selects functional roles, instantiates them as either single agents or reusable groups, and connects them to existing units, optimizing the construction policy with intermediate feedback from probe‑based utility and structural signals. Experiments show MAGIC outperforms state‑of‑the‑art baselines on eight benchmarks and achieves strong inference efficiency.

By Kairui Yang, Ziheng Yi, Xunkai Li, Minghao An, Zhanke Liu, Zekai Chen, Rong-Hua Li
arXiv AI
Aug 25

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.

By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv AI
Sep 12

DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat

The paper introduces DRG-MAPPO, a hierarchical multi‑agent reinforcement learning framework for cooperative air combat. It combines graph‑based relational modeling with dynamic role assignment, using a high‑level policy to allocate tactical roles such as leader and supporter, and a low‑level policy to execute maneuver actions. The approach includes a target‑priority auxiliary task and achieves an 87% win rate in experiments, indicating effective coordination and stability.

By Junlin Liu, Chengwei Li, Yang Gao, Hui Chang, Xinchen Zhang, Zhijun Zhao, Hao Zhao
arXiv AI
Sep 24

Reinforcement Learning with Decomposed Subtasks

The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that splits trajectory rewards into per‑subtask shares before policy updates, replacing the scalar advantage used in Group Relative Policy Optimization (GRPO). RLDS employs Subtask‑Decomposed Advantage Estimation (SDAE) to compute group‑relative advantages and distribute credit to tokens based on subtask importance, focusing on steps where a reflection marks a subtask as consequential. Experiments on four benchmarks—FrozenLake, HotpotQA, ScienceWorld, and DeepResearch—show that RLDS improves performance on high‑heterogeneity tasks (ScienceWorld and FrozenLake) and is more compute‑efficient than scalar GRPO for long rollouts.

By Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich