arXiv:2608. 14135v1 Announce Type: cross Abstract: Autonomous pursuit-evasion is a fundamental challenge for Unmanned Aerial Vehicles (UAVs), requiring rapid decision-making under tightly coupled dynamics and continuously changing opponent behaviors.
By Wenhao Tang, Tianyang Chen, Zhejun Cui, Boyuan An, Jiayu Chen, Ruize Zhang, Huidong Liu, Tianyue Wu, Qingmin Liao, Fei Gao, Yu Wang, Chao Yu
The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.
By Chuhao Qin, Evangelos Pournaras
arXiv:2607. 25728v1 Announce Type: cross Abstract: This paper presents a cooperative indoor UAV guidance framework that combines a shared voxel-map world model with a multi-agent Soft Actor-Critic (MASAC) controller.
By Thomas Hickling, Dylan Wynne, Yu Su, Nabil Aouf
arXiv:2603. 03741v2 Announce Type: replace-cross Abstract: To improve generalization and resilience in human-robot collaboration (HRC), robots must contend with diverse combinations of human behaviors and contexts, motivating multi-agent reinforcement learning (MARL).
By Hao Zhang, Yaru Niu, Yikai Wang, Ding Zhao, H. Eric Tseng
arXiv:2607. 21488v1 Announce Type: cross Abstract: Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) systems, which typically struggle with combinatorial action spaces, reliance on privileged information, or rigid agent designs.
By Gil Lifshits, Igal Bilik, Gilad Katz
AERIS is an offline policy improvement framework for multi-UAV integrated sensing and communication (ISAC) that learns from fixed flight logs using centralized training and decentralized execution. It introduces STAR-CRDT, an offline multi-agent RL algorithm that rectifies local actions and distills trusted improvements into decentralized actors, providing an offline-support policy improvement guarantee. Experiments demonstrate that STAR-CRDT boosts the main ISAC objective return by 29.3% and improves communication sum rate, sensing pass rate, and sensing margin while reducing collision-risk events by 54.2%.
By Ziyuan Wang (Steven), Yifan Sui (Steven), Wei Wei (Steven), Wenjie Xin (Steven), Zekai Zhang (Steven), Xiangwang Hou (Steven), Xiao-Ping (Steven), Zhang