ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
arXiv:2604. 13472v2 Announce Type: replace-cross Abstract: Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a centralized control problem into multiple interacting agents.
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
arXiv:2607. 19117v1 Announce Type: new Abstract: Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and continuous parameterization.
arXiv:2502. 00345v2 Announce Type: replace-cross Abstract: The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications.
arXiv:2606. 12281v1 Announce Type: cross Abstract: In Decentralized Training and Decentralized Execution (DTDE) for cooperative Multi-Agent Reinforcement Learning (MARL), action-advising-based knowledge sharing promotes interpretable and scalable cooperation among agents.
The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.
arXiv:2605.01457v4 Announce Type: replace Abstract: How can generative offline multi-agent reinforcement learning achieve both fast joint trajectory generation and effective cooperation? Multi-agent...
arXiv:2608. 11658v1 Announce Type: cross Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive.
The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.
arXiv:2610.01882v1 Announce Type: cross Abstract: Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment....
arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.
arXiv:2504. 16129v5 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based Multi-Agent Systems (LaMAS) have demonstrated strong capabilities on complex agentic tasks requiring multifaceted reasoning and collaboration, from high-quality presentation generation to scientific research.
Collab‑Solver introduces a multi‑agent policy learning framework for mixed‑integer linear programming (MILP) that enables collaborative optimization of multiple solver modules. By modeling the interaction between cut selection and branching as a Stackelberg game, the approach employs a two‑phase learning paradigm—data‑communicated policy pretraining followed by coordinated policy refinement. Experiments on synthetic and large‑scale real‑world MILP datasets show that the jointly learned policies markedly improve solving performance and generalize well across diverse instance sets.