Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs operate in a decentralized netwo...
The paper introduces a Physics‑Informed Multi‑Agent Coordination framework that embeds calibrated BCMP queueing topologies into a decentralized multi‑agent reinforcement learning system for hospital patient flow. It formulates the problem as a Decentralized Partially Observable Markov Decision Process with coupled resource constraints, enabling departmental agents to negotiate patient routing and service scaling while exchanging localized action fingerprints to handle non‑stationarity. Empirical tests on MIMIC‑IV data show the approach reduces cumulative system delay compared to static Markovian models, heuristic dispatching, and independent multi‑agent baselines, all while preserving clinical safety constraints.
By Guoqing Zhang, Rafik Hadfi, Takayuki Ito
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
By Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim
The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.
By Chuhao Qin, Evangelos Pournaras
The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.
By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv:2606. 11284v1 Announce Type: cross Abstract: Real-world multi-agent systems, from traffic coordination to resource allocation, are often modeled as general-sum games where individual incentives conflict with collective welfare.
By Wongyu Lee, Francesco Lelli, Omran Ayoub, Massimo Tornatore
arXiv:2607. 18359v1 Announce Type: cross Abstract: Critical infrastructures are increasingly distributed, interdependent, and exposed to evolving disruptions, making resilience a central requirement for their operation and control.
By Minghui Ding, Evangelos Pournaras
arXiv:2511. 13103v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations.
By Vidur Sinha, Muhammed Ustaomeroglu, Guannan Qu
arXiv:2604. 25848v2 Announce Type: replace Abstract: We study city-scale control of electric-vehicle (EV) ride-hailing fleets where dispatch, repositioning, and charging decisions must respect charger and feeder limits under uncertain, spatially correlated demand and travel times.
By An Nguyen, Hoang Nguyen, Phuong Le, Hung Pham, Cuong Do, Laurent El Ghaoui
arXiv:2503. 24183v3 Announce Type: replace Abstract: The expansion of ride-sourcing services such as Uber and Lyft has reshaped urban transportation by offering flexible, on-demand mobility via mobile applications.
By Matej Jusup, Kenan Zhang, Zhiyuan Hu, Barna P\'asztor, Andreas Krause, Francesco Corman
arXiv:2602. 20804v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) is typically framed as a decentralised partially observable Markov decision process (Dec-POMDP), a setting whose hardness stems from two key challenges: partial observability and decentralised coordination.
By Kale-ab Tessera, Leonard Hinckeldey, Riccardo Zamboni, David Abel, Amos Storkey
HySTAR is a MAPPO-based framework that addresses structural target drift in cooperative multi‑agent reinforcement learning by anchoring an overlapping sparse hypergraph as a stable high‑order value‑decomposition scaffold. It separates adaptive representation learning from a temporally consistent decomposition basis, using a spatiotemporal encoder to capture physical and task‑dependent interactions and combining temporal and structural relevance to compute agent‑specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE show consistent improvements over MAPPO‑style, value‑factorization, and dynamic‑grouping baselines, achieving significant gains in performance and convergence speed.
By Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang