arXiv AI

Symmetric Behavior Regularized Policy Optimization

arXiv:2508. 04225v4 Announce Type: replace-cross Abstract: Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline reinforcement learning.

arXiv Machine Learning
Jun 11

Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity

arXiv:2606. 11431v1 Announce Type: new Abstract: Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training.

By Shira Vansover-Hager, Matan Schliserman, Ofir Schlisselberg, Tomer Koren
arXiv Machine Learning
4d ago

Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning

The paper introduces Dually Regularized AIL, a model‑free algorithm for adversarial imitation learning that jointly applies KL policy regularization and a quadratic reward penalty based on expert and learner occupancies. It proves fast convergence rates, achieving a ×O(1/K+1/N) bound on the regularized imitation gap in finite‑horizon MDPs with general function approximation, and establishes the first algorithm to attain ×O(1/ε) sample complexity in both expert demonstrations and online interactions for this regularized objective.

By Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang
arXiv Machine Learning
Sep 14

A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning

The paper presents a unified framework for regularization-based robust reinforcement learning by deriving upper bounds on the performance gap between nominal and worst-case policies. These bounds are expressed as a regularization objective plus a KL-divergence penalty, explaining why KL penalties enhance robustness. The authors reformulate robust training as a constrained optimization problem, updating the Lagrange multiplier jointly with the policy to automatically tune regularization, and validate the approach with extensive adversarial evaluations on continuous control tasks.

By Amine Andam, Jamal Bentahar, Mustapha Hedabou