arXiv Machine Learning

Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning

arXiv:2507. 23604v2 Announce Type: replace Abstract: Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity.

arXiv Machine Learning
Aug 6

Communication-Enhanced Tutoring for Efficient Decentralized Multi-Agent Reinforcement Learning

arXiv:2508. 13661v4 Announce Type: replace Abstract: Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to act independently at test time while leveraging additional information during training.

By Maciej Wojtala, Bogusz Stefa\'nczyk, Dominik Bogucki, {\L}ukasz Lepak, Pawe{\l} Wawrzy\'nski
arXiv AI
Jul 1

HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation

arXiv:2606. 30966v1 Announce Type: new Abstract: Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives.

By Arshia Rafieioskouei, Tzu-Han Hsu, Matthew Lucas, Borzoo Bonakdarpour
arXiv AI
Jun 3

Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

arXiv:2602. 08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems.

By Yanming Li, Xuelin Zhang, WenJie Lu, Ziye Tang, Maodong Wu, Haotian Luo, Tongtong Wu, Zijie Peng, Hongze Mi, Yibo Feng, Naiqiang Tan, Chao Huang, Lian Peng, Li Shen
arXiv Machine Learning
Sep 24

Optimization without Future Compromises? Decentralized Coordination via Collective and Reinforcement Learning

The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.

By Chuhao Qin, Evangelos Pournaras
arXiv AI
Sep 12

Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning

The paper reviews hierarchical reinforcement learning (HRL) as a method for enabling agents to explore, plan, and learn in complex, open-ended environments by uncovering temporal structure in experience streams. It discusses the unclear definition of what makes a structure useful, the benefits of HRL for decision‑making challenges, and its impact on AI agent performance trade‑offs. The authors survey various HRL methods—from online learning to offline datasets and large language model integration—and outline the challenges and suitable domains for temporal structure discovery.

By Martin Klissarov, Akhil Bagaria, Ziyan Luo, George Konidaris, Doina Precup, Marlos C. Machado