arXiv Machine Learning

Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning

arXiv:2607. 03168v1 Announce Type: cross Abstract: Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation.

arXiv Machine Learning
5d ago

Robust Successor Features

The paper introduces robust successor features, a method that extends the successor representation to handle uncertainty in both reward functions and transition kernels within linear Markov Decision Processes. It provides a theoretical bound on Generalized Policy Improvement that quantifies performance loss due to mismatched dynamics, and demonstrates the approach on grid-based benchmarks against prior methods that consider only reward or transition differences.

By Erik Nikulski, Yamen Habib, Vicen\c{c} Gomez, Anders Jonsson, Rub\'en Moreno-Bote, Javier Segovia-Aguas
arXiv Machine Learning
Sep 14

A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning

The paper presents a unified framework for regularization-based robust reinforcement learning by deriving upper bounds on the performance gap between nominal and worst-case policies. These bounds are expressed as a regularization objective plus a KL-divergence penalty, explaining why KL penalties enhance robustness. The authors reformulate robust training as a constrained optimization problem, updating the Lagrange multiplier jointly with the policy to automatically tune regularization, and validate the approach with extensive adversarial evaluations on continuous control tasks.

By Amine Andam, Jamal Bentahar, Mustapha Hedabou
arXiv Machine Learning
Sep 15

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence

The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.

By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
arXiv Machine Learning
4d ago

Global Optimality for Constrained Exploration via Penalty Regularization

The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.

By Florian Wolf, Ilyas Fatkhullin, Niao He
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas