Bridging the Gap between Newton-Raphson Method and Regularized Policy Iteration
arXiv:2310. 07211v2 Announce Type: replace Abstract: Regularization is a cornerstone of modern reinforcement learning.
arXiv:2310. 07211v2 Announce Type: replace Abstract: Regularization is a cornerstone of modern reinforcement learning.
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.
arXiv:2512. 18336v2 Announce Type: replace-cross Abstract: This paper explores the impact of dynamic entropy tuning in Reinforcement Learning (RL) algorithms that train a stochastic policy.
arXiv:2607. 03168v1 Announce Type: cross Abstract: Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation.
arXiv:2610.02198v1 Announce Type: cross Abstract: Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic....
arXiv:2607. 22201v1 Announce Type: cross Abstract: We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions.
The paper introduces a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves the non-uniqueness of reward functions by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. It models expert demonstrations as a Markov chain, estimates the expert policy via penalized maximum likelihood, and provides high-probability bounds on the excess Kullback–Leibler divergence between the estimated and true policies. These results yield non-asymptotic minimax optimal convergence rates for the least-squares reward function, highlighting the trade-offs among smoothing, model complexity, and sample size.
arXiv:2606. 28671v1 Announce Type: new Abstract: Stackelberg differential games (SDGs) provide a powerful framework for hierarchical decision-making in stochastic and continuous-time environments, yet their solution remains computationally challenging due to the complexity of traditional dynamic programming and Hamilton-Jacobi-Bellman-Isaacs (HJBI) methods, especially in high-dimensional systems.
arXiv:2607. 23726v1 Announce Type: cross Abstract: Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning.
arXiv:2507.21543v3 Announce Type: replace-cross Abstract: Mutual information regularization has recently been studied in reinforcement learning (RL) as an extension of entropy regularization, in whic...
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.