arXiv:2508. 13661v4 Announce Type: replace Abstract: Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to act independently at test time while leveraging additional information during training.
By Maciej Wojtala, Bogusz Stefa\'nczyk, Dominik Bogucki, {\L}ukasz Lepak, Pawe{\l} Wawrzy\'nski
Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs operate in a decentralized netwo...
CoRe-MARL is a cooperative multi-agent reinforcement learning framework designed for decentralized relief distribution networks. It models each local center as an agent in a Dec-POMDP, using a recurrent network to learn redistribution policies that reduce service gaps and improve the worst-served region. Experiments show that recurrent MAPPO outperforms independent PPO and heuristic baselines, maintaining competitive network-wide service while adapting to evolving supply and demand dynamics.
By Naimur Rahman Chowdhury, Shatabdi Sen Prapti, Md. Salehin Seyam, Limon Bin Hossain
arXiv:2507. 10142v2 Announce Type: replace Abstract: Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the assumptions under which algorithms are designed and evaluated.
By Siyi Hu, Mohamad A Hady, Jianglin Qiao, Jimmy Cao, Mahardhika Pratama, Ryszard Kowalczyk
arXiv:2511. 13103v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations.
By Vidur Sinha, Muhammed Ustaomeroglu, Guannan Qu
The paper reviews the evolution of multi‑agent unmanned systems from isolated sensing to collaborative intelligence, where agents share compact features to overcome local observation limits such as occlusions and sensor range. It introduces a five‑dimensional taxonomy (collaboration stage, communication paradigm, fusion architecture, learning strategy, application domain) and three cognitive synergy conditions (Semantic Disambiguation, Pragmatic Information Exchange, Proactive Informational Foraging) to unify existing research. The authors survey architectures, neural‑communication co‑design, embodied action‑perception loops, and resilience mechanisms, map advances onto operational domains (V2X, UAV, logistics, smart cities), and propose the GCI‑Bench scoring protocol to standardize evaluation across studies.
By Lei Zhang, Chun Ye, Le Yang, Zhaozhong Wang, Deng-Ping Fan, Hang Dai, Binglu Wang