arXiv:2602. 04132v4 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has achieved remarkable success in solving complex sequential decision-making problems.
By Dhruv S. Kushwaha, Zoleikha A. Biron
arXiv:2606. 04749v1 Announce Type: cross Abstract: Safe robot control requires maximizing return while satisfying safety constraints.
By Guopeng Li, Moritz A. Zanger, Matthijs T. J. Spaan, Julian F. P. Kooij
arXiv:2608. 04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control.
By Mahshad Rastegarmoghaddam, Davoud Nikkhouy, Shima Samadzadeh
arXiv:2606. 14415v1 Announce Type: new Abstract: Safe reinforcement learning (Safe RL) aims to maximize expected return while satisfying safety constraints, typically modeled as Constrained Markov Decision Processes (CMDPs).
By Ayoub Belouadah, Sylvain Kubler, Yves Le Traon
arXiv:2606. 14536v1 Announce Type: new Abstract: Safe reinforcement learning (RL) aims to learn policies that optimize rewards while satisfying constraints.
By Kai S. Yun, Zeyang Li, Navid Azizan
arXiv:2501. 15373v2 Announce Type: replace-cross Abstract: Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance.
By Xinyang Wang, Hongwei Zhang, Shimin Wang, Wei Xiao, Martin Guay