arXiv:2602. 04132v4 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has achieved remarkable success in solving complex sequential decision-making problems.
By Dhruv S. Kushwaha, Zoleikha A. Biron
arXiv:2606. 04749v1 Announce Type: cross Abstract: Safe robot control requires maximizing return while satisfying safety constraints.
By Guopeng Li, Moritz A. Zanger, Matthijs T. J. Spaan, Julian F. P. Kooij
arXiv:2608. 04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control.
By Mahshad Rastegarmoghaddam, Davoud Nikkhouy, Shima Samadzadeh
The paper proves that using a permissive safety filter in reinforcement learning does not compromise asymptotic performance. By formalizing safety through a safety‑critical Markov decision process and a filtered MDP, the authors show that optimal policies in the filtered MDP achieve the same return as the best safe policy in the original setting. Experiments on Safety Gymnasium confirm zero violations during training and performance that matches or exceeds unfiltered baselines.
By Donggeon David Oh, Duy P. Nguyen, Haimin Hu, Jaime Fern\'andez Fisac
The paper introduces LEAP-CBF, a safety filter that uses Least‑Effort Adversarial Potentials to quantify how much disturbance effort is needed to cause failure in nonlinear dynamical systems. LEAP serves as a control barrier function for the undisturbed system and can be combined with a robust safety filter that tolerates disturbances with bounded cumulative effort. The authors develop a deep reinforcement learning method to construct LEAPs and validate their effectiveness through simulations of multi‑agent systems and hardware experiments on a quadruped and quadrotors.
By Oswin So, Eric Yu, Chuchu Fan
arXiv:2606. 14415v1 Announce Type: new Abstract: Safe reinforcement learning (Safe RL) aims to maximize expected return while satisfying safety constraints, typically modeled as Constrained Markov Decision Processes (CMDPs).
By Ayoub Belouadah, Sylvain Kubler, Yves Le Traon
arXiv:2606. 14536v1 Announce Type: new Abstract: Safe reinforcement learning (RL) aims to learn policies that optimize rewards while satisfying constraints.
By Kai S. Yun, Zeyang Li, Navid Azizan
arXiv:2501. 15373v2 Announce Type: replace-cross Abstract: Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance.
By Xinyang Wang, Hongwei Zhang, Shimin Wang, Wei Xiao, Martin Guay
arXiv:2607. 07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics.
By Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender
arXiv:2608. 10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints.
By Chenhua Fan, Jiahui Zhu, Yuhang Zhang, Honghao Wei
arXiv:2506. 02255v2 Announce Type: replace Abstract: Most existing safe reinforcement learning (RL) benchmarks focus on robotics and control tasks, offering limited relevance to high-stakes domains that involve structured constraints, mixed-integer decisions, and industrial complexity.
By Asha Ramanujam (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Adam Elyoumi (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Hao Chen (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Sai Madhukiran Kompalli (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Akshdeep Singh Ahluwalia (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Shraman Pal (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Dimitri J. Papageorgiou (Energy Sciences, ExxonMobil Technology and Engineering Company, Annandale, NJ), Can Li (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN)
arXiv:2609.13231v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provid...
By Manan Tayal, Akshay Nambi