Constrained Group Relative Policy Optimization
arXiv:2602.05863v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with const...
arXiv:2606. 02337v1 Announce Type: new Abstract: Constrained Multi-agent reinforcement learning (CMARL) faces two intertwined challenges: the joint action space grows exponentially with the number of agents, and additional requirements couple agents in ways that reward structure alone does not capture.
arXiv:2602.05863v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with const...
arXiv:2607. 21488v1 Announce Type: cross Abstract: Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) systems, which typically struggle with combinatorial action spaces, reliance on privileged information, or rigid agent designs.
Collab‑Solver introduces a multi‑agent policy learning framework for mixed‑integer linear programming (MILP) that enables collaborative optimization of multiple solver modules. By modeling the interaction between cut selection and branching as a Stackelberg game, the approach employs a two‑phase learning paradigm—data‑communicated policy pretraining followed by coordinated policy refinement. Experiments on synthetic and large‑scale real‑world MILP datasets show that the jointly learned policies markedly improve solving performance and generalize well across diverse instance sets.
arXiv:2608. 11658v1 Announce Type: cross Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive.
arXiv:2507. 23604v2 Announce Type: replace Abstract: Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity.
arXiv:2608. 05588v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones.
arXiv:2609.13739v1 Announce Type: cross Abstract: Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory...
arXiv:2609.06586v1 Announce Type: cross Abstract: A shared reward gives agents a common objective, but leaves open when, how and even whether they must cooperate to succeed. We address these question...
arXiv:2508. 14751v2 Announce Type: replace Abstract: We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces.
The paper introduces ControlG, a control‑theoretic framework for coordinating multi‑objective graph self‑supervised learning. It treats objective coordination as a temporal allocation problem, estimating each objective’s difficulty and antagonism, planning budgets with a Pareto‑aware log‑hypervolume planner, and scheduling updates via a PID controller. Experiments on nine datasets show that ControlG consistently outperforms state‑of‑the‑art baselines and provides an auditable schedule revealing which objectives drive learning.
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.