arXiv AI By Jose Tupayachi, Xueping Li, Soham Das

Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework

Read the original on arXiv AI →

The paper introduces a Lagrangian framework for managing a regenerative commons, framing the problem as a constrained Markov game with a specified depletion budget. It constructs policy sequences from unconstrained solutions, extending time‑average concepts to reset episodes with discounted rewards and terminal costs, and provides theoretical guarantees such as reward‑independent feasibility, cooperative feasibility, and approximate optimality. Experiments on a fishery model using constrained IPPO and MAPPO illustrate how depletion budgets influence stock retention, harvest rewards, and price adaptation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning

The paper introduces Exchange Policy Optimization (EPO), a framework for semi‑infinite safe reinforcement learning that handles infinitely many constraints by iteratively solving finite subproblems. EPO expands or deletes constraints based on tolerance violations and Lagrange multipliers, maintaining computational tractability while converging to an optimal policy with bounded safety violations. The authors prove finite convergence, provide iteration bounds, and quantify the suboptimality gap under mild assumptions.

By Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li
arXiv Machine Learning
2d ago

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

The paper introduces a new approach to learning chance-constrained Markov decision processes (CCMDPs) using a Bellman distributional certificate. It provides both model-based and model-free algorithms with theoretical guarantees, including matching upper and lower bounds for tabular discounted CCMDPs with bounded successor support. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage benchmark demonstrate the safety and effectiveness of the proposed methods.

By Chenbei Lu, Hongyu Yi
arXiv Machine Learning
Jul 30

Stable and Budget-Feasible Coalition Formation for Clustered Federated Learning: A Hedonic Potential-Game Approach

arXiv:2607. 26788v1 Announce Type: cross Abstract: Clustered federated learning benefits from organizing heterogeneous participants into coalitions that train coalition-specific models, but such clustering is sustainable only if participants prefer their assigned coalition and the required transfers are affordable.

By Cengis Hasan