arXiv AI

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning

arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.

arXiv Machine Learning
Aug 6

Communication-Enhanced Tutoring for Efficient Decentralized Multi-Agent Reinforcement Learning

arXiv:2508. 13661v4 Announce Type: replace Abstract: Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to act independently at test time while leveraging additional information during training.

By Maciej Wojtala, Bogusz Stefa\'nczyk, Dominik Bogucki, {\L}ukasz Lepak, Pawe{\l} Wawrzy\'nski
arXiv AI
Sep 17

CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning

CoRe-MARL is a cooperative multi-agent reinforcement learning framework designed for decentralized relief distribution networks. It models each local center as an agent in a Dec-POMDP, using a recurrent network to learn redistribution policies that reduce service gaps and improve the worst-served region. Experiments show that recurrent MAPPO outperforms independent PPO and heuristic baselines, maintaining competitive network-wide service while adapting to evolving supply and demand dynamics.

By Naimur Rahman Chowdhury, Shatabdi Sen Prapti, Md. Salehin Seyam, Limon Bin Hossain
arXiv Machine Learning
Sep 24

Optimization without Future Compromises? Decentralized Coordination via Collective and Reinforcement Learning

The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.

By Chuhao Qin, Evangelos Pournaras
arXiv AI
Sep 28

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.

By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu