arXiv AI By Ayoub Belouadah, Sylvain Kubler, Yves Le Traon

CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning

Read the original on arXiv AI →

arXiv:2606. 14415v1 Announce Type: new Abstract: Safe reinforcement learning (Safe RL) aims to maximize expected return while satisfying safety constraints, typically modeled as Constrained Markov Decision Processes (CMDPs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning

The paper introduces Exchange Policy Optimization (EPO), a framework for semi‑infinite safe reinforcement learning that handles infinitely many constraints by iteratively solving finite subproblems. EPO expands or deletes constraints based on tolerance violations and Lagrange multipliers, maintaining computational tractability while converging to an optimal policy with bounded safety violations. The authors prove finite convergence, provide iteration bounds, and quantify the suboptimality gap under mild assumptions.

By Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li