Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback.
arXiv:2606. 24622v1 Announce Type: new Abstract: Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors.
By Andreas Chouliaras, Luke Connolly, Dimitris Chatzpoulos
We’re releasing Safety Gym, a suite of environments and tools for measuring progress towards reinforcement learning agents that respect safety constraints while training.
We’re releasing the public beta of OpenAI Gym, a toolkit for developing and comparing reinforcement learning (RL) algorithms. It consists of a growing suite of environments (from simulated robots to Atari games), and a site for comparing and reproducing results.
arXiv:1908.08773v3 Announce Type: replace
Abstract: In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit...
By Victor Gallego, Roi Naveiro, David Rios Insua, David Gomez-Ullate Oteiza
arXiv:2607. 18314v1 Announce Type: new Abstract: Experiment trackers show how training is progressing, but changing a live run still usually requires trainer-specific code.
By Wentao Zhang, Xuanhe Pan, Han Zhou, Yang Lu, Yuntian Deng
arXiv:2607. 05773v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making.
By Akshay Arora, Ishan Nigam, Ashutosh Aggarwal, Shefali Bansal, Krishna Singh, Sweta Kumari, Nikhil Mittal, Shariq Farhan, Siddarth Malreddy
The paper presents an algorithm that lets a learning agent ask for help from a mentor and transfer knowledge between similar states, enabling safe and effective learning in Markov decision processes with irreversible dynamics and infinite state spaces. It proves that both regret and the number of mentor queries grow sublinearly over time, using a sequence of three reductions to achieve a general result. The work claims to be the first formal proof that an agent can achieve high reward while becoming self‑sufficient in an unknown, unbounded, high‑stakes environment without resets.
By Benjamin Plaut, Juan Li\'evano-Karim, Hanlin Zhu, Stuart Russell
arXiv:2606. 09559v1 Announce Type: cross Abstract: Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such as robotics systems.
By Shixiong Jiang, Taozheng Zhu, Fanxin Kong
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can b...
arXiv:2607. 18488v1 Announce Type: cross Abstract: Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology remains rooted in simulations.
By Elena Sorina Lupu, Patrick Spieler, Khurram Javed, Kris De Asis, John D. Martin, Martha Steenstrup, Joseph Modayil