arXiv Machine Learning By Shreeram Murali, Shankar A. Deka, Dominik Baumann

Computationally efficient safe exploration in reinforcement learning

Read the original on arXiv Machine Learning →

The paper introduces “CoLSafe-MDP”, a reinforcement learning algorithm that ensures safe exploration in constrained Markov decision processes. It replaces computationally heavy Gaussian process methods with a Nadaraya-Watson estimator, achieving constant-time scaling for estimate bounds. The authors evaluate the algorithm on a grid-based environment and on observational Martian terrain data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 17

Expected Free Energy-based Informative Path Planning for Robotic Mars Exploration

arXiv:2608. 14466v1 Announce Type: cross Abstract: An autonomous robot efficiently exploring an unknown environment, such as looking for water sources on Mars, faces two simultaneous demands: building an accurate information map while quickly finding the regions of greatest value, and paying for every meter of travel and the cost of every measurement it takes.

By Ajith Anil Meera, Pablo Lanillos, Wouter Kouw
arXiv AI
Aug 21

Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning

arXiv:2608. 19836v1 Announce Type: cross Abstract: Probabilistic shielding is a technique for safe reinforcement learning (RL).

By Astrid Horn Brorholt (Aalborg University, Aalborg, Denmark), Maris F. L. Galesloot (Radboud University, Nijmegen, Netherlands), Nils Jansen (Radboud University, Nijmegen, Netherlands), Kim Guldstrand Larsen (Aalborg University, Aalborg, Denmark), Christian Schilling (Aalborg University, Aalborg, Denmark)
arXiv Machine Learning
Sep 18

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

The paper introduces a model-based bootstrap framework for uncertainty quantification in offline policy evaluation (OPE) within finite-horizon, time-inhomogeneous Markov decision processes. Unlike traditional bootstrap methods that resample entire episodes, this approach regenerates trajectories from an estimated MDP, enabling use of diverse offline data formats such as complete trajectories, transition-level observations, and trajectory fragments. The authors prove bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation, and demonstrate through simulations that the method yields tighter confidence intervals and more accurate variance estimates compared to existing techniques.

By Weiwei Wang, Yuqiang Li, Xianyi Wu, Bingyi Jing
arXiv AI
Jul 3

Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

arXiv:2607. 01794v1 Announce Type: cross Abstract: With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection, environmental monitoring, and rescue, creating growing demand for reliable autonomous navigation.

By Shenghui Zhang, YuXuan Gao, Songwei Zhao, Jifeng Hu, Zijing Zhang, Hechang Chen