The paper presents an algorithm that lets a learning agent ask for help from a mentor and transfer knowledge between similar states, enabling safe and effective learning in Markov decision processes with irreversible dynamics and infinite state spaces. It proves that both regret and the number of mentor queries grow sublinearly over time, using a sequence of three reductions to achieve a general result. The work claims to be the first formal proof that an agent can achieve high reward while becoming self‑sufficient in an unknown, unbounded, high‑stakes environment without resets.
By Benjamin Plaut, Juan Li\'evano-Karim, Hanlin Zhu, Stuart Russell
arXiv:2606. 24991v1 Announce Type: cross Abstract: Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning.
By Shambhuraj Sawant, Akhil S Anand, Dirk Reinhardt, Sebastien Gros
arXiv:2606. 00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs).
By Edwin Hamel-De le Court, Thom Badings, Alessandro Abate, Francesco Belardinelli, Francesco Fabiano
arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.
By Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning. However, despite these strengths, an MPC scheme typically does not yield optimal policies for sequential decision-making problems formulated as Markov Decision Processes (MDPs).
The paper introduces Quasar, a model‑free Q‑learning algorithm that guarantees asymptotic convergence for reachability objectives in Markov Decision Processes that are free of non‑terminal maximal end components (MECs). Unlike prior model‑based methods, Quasar does not estimate transition probabilities, reducing memory usage from O(|S|²|A|) to O(|S||A|). Experiments on the Quantitative Verification Benchmark Set show that Quasar converges to optimal policies with far fewer samples than existing state‑of‑the‑art model‑based approaches.
By Lu-Chin Chang, Suguman Bansal