arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.
By Volodymyr Tkachuk, Csaba Szepesv\'ari, Xiaoqi Tan
arXiv:2406.03894v2 Announce Type: replace
Abstract: Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due...
By Yaozhong Gan, Renye Yan, Xiaoyang Tan, Zhe Wu, Junliang Xing
arXiv:2406.03678v2 Announce Type: replace
Abstract: On-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensi...
By Yaozhong Gan, Renye Yan, Zhe Wu, Junliang Xing
arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.
By Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
The paper offers a new way to view the occupancy measure in reinforcement learning by embedding the planning criterion into the dynamics via a resetting planning process. The resulting stationary measure, called the visitation measure, forms a dually flat statistical manifold with two affine charts: visitation probabilities and log-policies, which are dual under conditional entropy. This geometric framework allows planning-as-inference to extend beyond linear rewards to nonlinear functionals of visitation, with each iteration solvable by a natural-gradient step and provides a new interpretation of the temporal-difference error as a marginal-utility estimate.
By Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs, Kenji Doya, Nico Scherf