arXiv:2608. 19836v1 Announce Type: cross Abstract: Probabilistic shielding is a technique for safe reinforcement learning (RL).
By Astrid Horn Brorholt (Aalborg University, Aalborg, Denmark), Maris F. L. Galesloot (Radboud University, Nijmegen, Netherlands), Nils Jansen (Radboud University, Nijmegen, Netherlands), Kim Guldstrand Larsen (Aalborg University, Aalborg, Denmark), Christian Schilling (Aalborg University, Aalborg, Denmark)
The paper offers a new way to view the occupancy measure in reinforcement learning by embedding the planning criterion into the dynamics via a resetting planning process. The resulting stationary measure, called the visitation measure, forms a dually flat statistical manifold with two affine charts: visitation probabilities and log-policies, which are dual under conditional entropy. This geometric framework allows planning-as-inference to extend beyond linear rewards to nonlinear functionals of visitation, with each iteration solvable by a natural-gradient step and provides a new interpretation of the temporal-difference error as a marginal-utility estimate.
By Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs, Kenji Doya, Nico Scherf
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.
By Volodymyr Tkachuk, Csaba Szepesv\'ari, Xiaoqi Tan
The paper introduces a framework for combining large language models (LLMs) with reinforcement learning (RL) by treating the LLM as a planner and the RL agent as a controller. It formalizes this hybrid setup as a Goal-Augmented Markov Decision Process and proves that using the LLM’s per‑state progress score as a bounded potential function preserves the optimal policy set, even if the LLM scores are inaccurate. The authors validate their theoretical result with numerical experiments on a small MDP, testing four potential configurations, including an adversarial case with a potential scaled twenty times the base reward.
By Christophe D. Hounwanou, John Emeka Eze, Ya\'e U. Gaba