arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
By Zhaoyu Zhu, Rui Gao, Shuang Li
arXiv:2608.22636v1 Announce Type: cross
Abstract: Q-learning with linear function approximation can be unstable because an arbitrary approximation architecture need not preserve the Bellman contracti...
By Shengbo Wang
The paper introduces a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves the non-uniqueness of reward functions by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. It models expert demonstrations as a Markov chain, estimates the expert policy via penalized maximum likelihood, and provides high-probability bounds on the excess Kullback–Leibler divergence between the estimated and true policies. These results yield non-asymptotic minimax optimal convergence rates for the least-squares reward function, highlighting the trade-offs among smoothing, model complexity, and sample size.
By Denis Belomestny, Alexey Naumov, Artemy Rubtsov, Sergey Samsonov
arXiv:2606. 16759v1 Announce Type: new Abstract: We study inverse reinforcement learning for discrete-time, infinite-horizon mean-field games (MFGs) under an average-reward criterion.
By \c{S}evket Kaan Alk{\i}r, Naci Sald{\i}, Berkay Anahtarc{\i}, Can Deha Kar{\i}ks{\i}z
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first st...
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class.
arXiv:2609.24103v1 Announce Type: new
Abstract: In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. W...
By Larry Preuett, Qiuyi Zhang, Muhammad Aurangzeb Ahmad
arXiv:2606. 16515v1 Announce Type: cross Abstract: Hamilton-Jacobi-Bellman theory implies that the optimal goal-conditioned action depends on the goal only through the gradient of the goal-reaching distance at the current state, yet standard online GCRL still conditions the actor on the raw goal -- a signal that is geometrically uninformative when the goal is far from the data distribution.
By Swaminathan S K, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra
arXiv:2310. 07211v2 Announce Type: replace Abstract: Regularization is a cornerstone of modern reinforcement learning.
By Zeyang Li, Chuxiong Hu, Yunan Wang, Guojian Zhan, Jie Li, Yao Lyu, Shengbo Eben Li
arXiv:2607. 05375v1 Announce Type: cross Abstract: Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation.
By Lars van der Laan, Nathan Kallus
The paper introduces a new approach to learning chance-constrained Markov decision processes (CCMDPs) using a Bellman distributional certificate. It provides both model-based and model-free algorithms with theoretical guarantees, including matching upper and lower bounds for tabular discounted CCMDPs with bounded successor support. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage benchmark demonstrate the safety and effectiveness of the proposed methods.
By Chenbei Lu, Hongyu Yi
arXiv:2605. 08253v2 Announce Type: replace Abstract: Distributional reinforcement learning (DRL) models the full return distribution, but existing finite-support or quantile-based methods rely on projections, while recent flow-based approaches can suffer from \emph{boundary mismatch} at the flow source or from \emph{high-variance} bootstrapping when current and successor noises are independent.
By Boyang Xu, Qing Zou, Siqin Yang, Hao Yan