arXiv Machine Learning By Zijian Zhao, Sen Li

Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

Read the original on arXiv Machine Learning →

arXiv:2607. 17924v1 Announce Type: cross Abstract: Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

HySTAR is a MAPPO-based framework that addresses structural target drift in cooperative multi‑agent reinforcement learning by anchoring an overlapping sparse hypergraph as a stable high‑order value‑decomposition scaffold. It separates adaptive representation learning from a temporally consistent decomposition basis, using a spatiotemporal encoder to capture physical and task‑dependent interactions and combining temporal and structural relevance to compute agent‑specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE show consistent improvements over MAPPO‑style, value‑factorization, and dynamic‑grouping baselines, achieving significant gains in performance and convergence speed.

By Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang
arXiv Machine Learning
Jul 30

Stable and Budget-Feasible Coalition Formation for Clustered Federated Learning: A Hedonic Potential-Game Approach

arXiv:2607. 26788v1 Announce Type: cross Abstract: Clustered federated learning benefits from organizing heterogeneous participants into coalitions that train coalition-specific models, but such clustering is sustainable only if participants prefer their assigned coalition and the required transfers are affordable.

By Cengis Hasan
arXiv AI
Jun 9

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

arXiv:2603. 21563v4 Announce Type: replace Abstract: Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such systems is limited by credit assignment: shared terminal rewards obscure individual contributions and can encourage free-riding.

By Zhongyi Li, Wan Tian, Yikun Ban, Jinju Chen, Huiming Zhang, Yang Liu, Fuzhen Zhuang