Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
The review explores how control theory, optimal transport, probabilistic inference, non‑equilibrium thermodynamics, and machine learning are interconnected through the optimization of free‑energy‑like functionals under dynamical or statistical constraints. It presents a conceptual thread linking these five fields and illustrates the ideas with applications in reinforcement learning, variational inference, and generative modeling. The article is written for readers without prior familiarity, beginning with physics principles.
arXiv:2102. 09235v3 Announce Type: replace Abstract: Recent studies revealed the mathematical connection between deep neural networks (DNNs) and dynamic systems.
arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.
The paper develops a diffusion approximation for stochastic gradient descent (SGD) when the optimization target is a functional on the Wasserstein space ℝ2. By lifting the problem to a Hilbert space via Lions differentiability, the authors construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. They prove that this Gaussian approximation achieves second‑order weak accuracy, providing a rigorous basis for replacing sample‑driven randomness with analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.
arXiv:2506. 04480v2 Announce Type: replace-cross Abstract: This paper focuses on Geodesic Principal Component Analysis (GPCA) on a collection of probability distributions using the Otto-Wasserstein geometry.