arXiv:2502. 00345v2 Announce Type: replace-cross Abstract: The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications.
By Yurui Li, Yuxuan Chen, Xiaoli Yang, Shijian Li, Gang Pan
arXiv:2508. 13661v4 Announce Type: replace Abstract: Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to act independently at test time while leveraging additional information during training.
By Maciej Wojtala, Bogusz Stefa\'nczyk, Dominik Bogucki, {\L}ukasz Lepak, Pawe{\l} Wawrzy\'nski
arXiv:2507. 23604v2 Announce Type: replace Abstract: Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity.
By Tommaso Marzi, Cesare Alippi, Andrea Cini
arXiv:2607. 18719v1 Announce Type: cross Abstract: This study proposes a learning method for multi-agent systems that allows agents to be controlled through human manager instructions after learning and enables uninstructed agents to implicitly complement the overall work based on the actions of other agents.
By Yamato Takahagi, Gentoku Nakasone, Yoshinari Motokawa, Toshiharu Sugawara
arXiv:2608. 08604v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) is a powerful framework for solving complex collaborative tasks, but it relies heavily on well-defined global reward functions.
By Ni Mu, Yao Luan, Yiqin Yang, Qing-Shan Jia
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
By Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim
arXiv:2605. 12655v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) in real-world use cases may need to adapt to external natural language instructions that interrupt ongoing behavior and conflict with long-horizon objectives.
By Wo Wei Lin, Ethan Rathbun, Enrico Marchesini, Xiang Zhi Tan
arXiv:2606. 14693v1 Announce Type: cross Abstract: Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting objectives.
By Pengxin Wang, Lihao Guo, Yi Xie, Bo Liu, Siyang Cao, Jingdi Chen
The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.
By Chuhao Qin, Evangelos Pournaras
arXiv:2605. 18077v2 Announce Type: replace Abstract: Communication is a key component in multi-agent reinforcement learning (MARL) for mitigating partial observability, yet prior approaches often rely on inefficient information exchange or fail to transmit sufficient state information.
By Sangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul Han
arXiv:2609.06586v1 Announce Type: cross
Abstract: A shared reward gives agents a common objective, but leaves open when, how and even whether they must cooperate to succeed. We address these question...
By Yannick Molinghen, Hugo Charels, Tom Lenaerts
Collab‑Solver introduces a multi‑agent policy learning framework for mixed‑integer linear programming (MILP) that enables collaborative optimization of multiple solver modules. By modeling the interaction between cut selection and branching as a Stackelberg game, the approach employs a two‑phase learning paradigm—data‑communicated policy pretraining followed by coordinated policy refinement. Experiments on synthetic and large‑scale real‑world MILP datasets show that the jointly learned policies markedly improve solving performance and generalize well across diverse instance sets.
By Siyuan Li, Yifan Yu, Zhihao Zhang, Mengjing Chen, Fangzhou Zhu, Tao Zhong, Peng Liu, Jianye Hao