Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning
arXiv:2608. 19836v1 Announce Type: cross Abstract: Probabilistic shielding is a technique for safe reinforcement learning (RL).
The paper discusses how reinforcement learning theory relies on probability theory via Markov chains and highlights a deep link between probability theory and potential theory. It reviews this connection and examines how a potential-theoretic perspective can be applied to core RL representations and algorithms under a fixed‑policy assumption, suggesting possible gains in sample efficiency and formal constraints. The authors also note that relaxing the fixed‑policy assumption allows the linear potential theory framework to extend naturally to nonlinear cases.
arXiv:2608. 19836v1 Announce Type: cross Abstract: Probabilistic shielding is a technique for safe reinforcement learning (RL).
The paper offers a new way to view the occupancy measure in reinforcement learning by embedding the planning criterion into the dynamics via a resetting planning process. The resulting stationary measure, called the visitation measure, forms a dually flat statistical manifold with two affine charts: visitation probabilities and log-policies, which are dual under conditional entropy. This geometric framework allows planning-as-inference to extend beyond linear rewards to nonlinear functionals of visitation, with each iteration solvable by a natural-gradient step and provides a new interpretation of the temporal-difference error as a marginal-utility estimate.
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.
The paper introduces a framework for combining large language models (LLMs) with reinforcement learning (RL) by treating the LLM as a planner and the RL agent as a controller. It formalizes this hybrid setup as a Goal-Augmented Markov Decision Process and proves that using the LLM’s per‑state progress score as a bounded potential function preserves the optimal policy set, even if the LLM scores are inaccurate. The authors validate their theoretical result with numerical experiments on a small MDP, testing four potential configurations, including an adversarial case with a potential scaled twenty times the base reward.
arXiv:2511. 03618v2 Announce Type: replace Abstract: In this paper, we formalize the almost sure convergence of $Q$-learning and linear temporal difference (TD) learning with Markovian samples using the Lean 4 theorem prover based on the Mathlib library.
arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.
arXiv:2307. 10524v3 Announce Type: replace Abstract: We study the tradeoff between consistency and robustness in the context of a single-trajectory time-varying Markov Decision Process (MDP) with untrusted machine-learned advice.
arXiv:2607. 20152v1 Announce Type: cross Abstract: Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle.
arXiv:2607. 03168v1 Announce Type: cross Abstract: Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation.
arXiv:2606. 24991v1 Announce Type: cross Abstract: Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning.