arXiv Machine Learning

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning

arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.

arXiv Machine Learning
4d ago

Global Optimality for Constrained Exploration via Penalty Regularization

The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.

By Florian Wolf, Ilyas Fatkhullin, Niao He
arXiv Machine Learning
Jul 21

Distributional Soft Bellman Operator under the Cram\'er Geometry

arXiv:2607. 17897v1 Announce Type: new Abstract: Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns.

By Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li
arXiv Machine Learning
Sep 17

Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization

The paper introduces a geometric framework for reinforcement learning that treats policies as mappings into the Wasserstein space of action probabilities. It establishes a Riemannian structure induced by stationary distributions, defines the tangent space of policies, and characterizes geodesics while addressing measurability concerns. The authors formulate a general RL optimization problem, construct a gradient flow via Otto's calculus, compute the gradient and Hessian of the energy, and demonstrate the approach with numerical examples for low‑dimensional problems and neural‑network‑parameterized policies for high‑dimensional settings.

By Mathias Dus (IRMA)
arXiv Machine Learning
Jul 22

Linear convergence of proximal descent schemes on the Wasserstein space

arXiv:2411. 15067v2 Announce Type: replace-cross Abstract: We investigate proximal descent methods, inspired by the minimizing movement scheme introduced by Jordan, Kinderlehrer and Otto, for optimizing entropy-regularized functionals on the Wasserstein space.

By Razvan-Andrei Lascu, Mateusz B. Majka, David \v{S}i\v{s}ka, {\L}ukasz Szpruch
arXiv Machine Learning
Sep 15

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence

The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.

By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
arXiv AI
Jun 16

Direction-Conditioned Policies via Compositional Subgoal Scoring for Online Goal-Conditioned Reinforcement Learning

arXiv:2606. 16515v1 Announce Type: cross Abstract: Hamilton-Jacobi-Bellman theory implies that the optimal goal-conditioned action depends on the goal only through the gradient of the goal-reaching distance at the current state, yet standard online GCRL still conditions the actor on the raw goal -- a signal that is geometrically uninformative when the goal is far from the data distribution.

By Swaminathan S K, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas