arXiv Machine Learning

Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

arXiv:2607. 17924v1 Announce Type: cross Abstract: Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL).

arXiv Machine Learning
5d ago

HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

HySTAR is a MAPPO-based framework that addresses structural target drift in cooperative multi‑agent reinforcement learning by anchoring an overlapping sparse hypergraph as a stable high‑order value‑decomposition scaffold. It separates adaptive representation learning from a temporally consistent decomposition basis, using a spatiotemporal encoder to capture physical and task‑dependent interactions and combining temporal and structural relevance to compute agent‑specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE show consistent improvements over MAPPO‑style, value‑factorization, and dynamic‑grouping baselines, achieving significant gains in performance and convergence speed.

By Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang
arXiv Machine Learning
Jul 30

Stable and Budget-Feasible Coalition Formation for Clustered Federated Learning: A Hedonic Potential-Game Approach

arXiv:2607. 26788v1 Announce Type: cross Abstract: Clustered federated learning benefits from organizing heterogeneous participants into coalitions that train coalition-specific models, but such clustering is sustainable only if participants prefer their assigned coalition and the required transfers are affordable.

By Cengis Hasan
arXiv AI
Jun 9

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

arXiv:2603. 21563v4 Announce Type: replace Abstract: Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such systems is limited by credit assignment: shared terminal rewards obscure individual contributions and can encourage free-riding.

By Zhongyi Li, Wan Tian, Yikun Ban, Jinju Chen, Huiming Zhang, Yang Liu, Fuzhen Zhuang
arXiv Machine Learning
Jul 27

Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning

arXiv:2601. 17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observability.

By Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif
arXiv AI
Sep 24

Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence

The paper investigates whether evolutionary stability guarantees that learning agents can achieve cooperative outcomes in a multi‑agent setting. Using a three‑agent governance game, the authors compare the evolutionary basin of attraction with learning basins derived from independent Q‑learning, scaled Boltzmann exploration, and SA–EA BQL. They find that while the evolutionary basin covers the entire sampled grid, only ε‑greedy Q‑learning attains a substantial learning basin, whereas the other methods fail to sustain cooperation, highlighting a disconnect between population‑level stability and finite‑sample learning accessibility.

By Yijie Wang