arXiv Machine Learning

PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks

arXiv:2607. 17922v1 Announce Type: cross Abstract: Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state unexamined.

Hugging Face Trending Papers
Jul 20

PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks

Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state unexamined. We show that sustained non-stationarity damages this internal state directly: as objectives shift, neurons progressively fall dormant and the shared policy loses the capacity to learn.

arXiv AI
Jul 7

Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

arXiv:2607. 04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood.

By Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi
arXiv Machine Learning
Jul 21

Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss

arXiv:2607. 17914v1 Announce Type: cross Abstract: Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments.

By Kemal Devrim Kafadar, Eren \"Ozaltun, Mahmud Efnan \c{S}anl{\i}, Feyza Orak, Emirhan Gazi, Kubilay Ka\u{g}an K\"om\"urc\"u, Naz{\i}m Kemal \"Ure
arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv Machine Learning
Aug 7

Communication-Aware Multi-Agent Reinforcement Learning for Decentralized Cooperative UAV Deployment

arXiv:2603. 16141v2 Announce Type: replace-cross Abstract: Autonomous Unmanned Aerial Vehicle (UAV) swarms are increasingly used as rapidly deployable aerial relays and sensing platforms, yet practical deployments must operate under partial observability and intermittent peer-to-peer connectivity.

By Enguang Fan, Yifan Chen, Zihan Shan, Klara Nahrstedt, Matthew Caesar, Jae Kim
arXiv AI
Aug 25

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.

By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv Machine Learning
Aug 27

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication

AERIS is an offline policy improvement framework for multi-UAV integrated sensing and communication (ISAC) that learns from fixed flight logs using centralized training and decentralized execution. It introduces STAR-CRDT, an offline multi-agent RL algorithm that rectifies local actions and distills trusted improvements into decentralized actors, providing an offline-support policy improvement guarantee. Experiments demonstrate that STAR-CRDT boosts the main ISAC objective return by 29.3% and improves communication sum rate, sensing pass rate, and sensing margin while reducing collision-risk events by 54.2%.

By Ziyuan Wang (Steven), Yifan Sui (Steven), Wei Wei (Steven), Wenjie Xin (Steven), Zekai Zhang (Steven), Xiangwang Hou (Steven), Xiao-Ping (Steven), Zhang