arXiv AI

Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning

arXiv AI
Sep 30

Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework

The paper introduces a Lagrangian framework for managing a regenerative commons, framing the problem as a constrained Markov game with a specified depletion budget. It constructs policy sequences from unconstrained solutions, extending time‑average concepts to reset episodes with discounted rewards and terminal costs, and provides theoretical guarantees such as reward‑independent feasibility, cooperative feasibility, and approximate optimality. Experiments on a fishery model using constrained IPPO and MAPPO illustrate how depletion budgets influence stock retention, harvest rewards, and price adaptation.

By Jose Tupayachi, Xueping Li, Soham Das
arXiv Machine Learning
Jul 30

Stable and Budget-Feasible Coalition Formation for Clustered Federated Learning: A Hedonic Potential-Game Approach

arXiv:2607. 26788v1 Announce Type: cross Abstract: Clustered federated learning benefits from organizing heterogeneous participants into coalitions that train coalition-specific models, but such clustering is sustainable only if participants prefer their assigned coalition and the required transfers are affordable.

By Cengis Hasan
arXiv AI
Jun 2

Coordination Graphs for Constrained Multi-Agent Reinforcement Learning

arXiv:2606. 02337v1 Announce Type: new Abstract: Constrained Multi-agent reinforcement learning (CMARL) faces two intertwined challenges: the joint action space grows exponentially with the number of agents, and additional requirements couple agents in ways that reward structure alone does not capture.

By Santiago Amaya-Corredor, Miguel Calvo-Fullana, Anders Jonsson
arXiv AI
Jun 3

Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

arXiv:2602. 08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems.

By Yanming Li, Xuelin Zhang, WenJie Lu, Ziye Tang, Maodong Wu, Haotian Luo, Tongtong Wu, Zijie Peng, Hongze Mi, Yibo Feng, Naiqiang Tan, Chao Huang, Lian Peng, Li Shen
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen