arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
By Zhaoyu Zhu, Rui Gao, Shuang Li
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
The paper introduces a geometric framework for reinforcement learning that treats policies as mappings into the Wasserstein space of action probabilities. It establishes a Riemannian structure induced by stationary distributions, defines the tangent space of policies, and characterizes geodesics while addressing measurability concerns. The authors formulate a general RL optimization problem, construct a gradient flow via Otto's calculus, compute the gradient and Hessian of the energy, and demonstrate the approach with numerical examples for low‑dimensional problems and neural‑network‑parameterized policies for high‑dimensional settings.
By Mathias Dus (IRMA)
arXiv:2607. 22201v1 Announce Type: cross Abstract: We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions.
By Mintae Kim, Koushil Sreenath
arXiv:2606. 16759v1 Announce Type: new Abstract: We study inverse reinforcement learning for discrete-time, infinite-horizon mean-field games (MFGs) under an average-reward criterion.
By \c{S}evket Kaan Alk{\i}r, Naci Sald{\i}, Berkay Anahtarc{\i}, Can Deha Kar{\i}ks{\i}z
arXiv:2607. 11005v1 Announce Type: cross Abstract: This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the dynamics and reward may depend on the joint distribution of states and controls.
By Ziheng Cheng, Xin Guo, Huy\^en Pham, Yufei Zhang