arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.
By Volodymyr Tkachuk, Csaba Szepesv\'ari, Xiaoqi Tan
arXiv:2406.03894v2 Announce Type: replace
Abstract: Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due...
By Yaozhong Gan, Renye Yan, Xiaoyang Tan, Zhe Wu, Junliang Xing
arXiv:2406.03678v2 Announce Type: replace
Abstract: On-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensi...
By Yaozhong Gan, Renye Yan, Zhe Wu, Junliang Xing
arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.
By Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
The paper offers a new way to view the occupancy measure in reinforcement learning by embedding the planning criterion into the dynamics via a resetting planning process. The resulting stationary measure, called the visitation measure, forms a dually flat statistical manifold with two affine charts: visitation probabilities and log-policies, which are dual under conditional entropy. This geometric framework allows planning-as-inference to extend beyond linear rewards to nonlinear functionals of visitation, with each iteration solvable by a natural-gradient step and provides a new interpretation of the temporal-difference error as a marginal-utility estimate.
By Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs, Kenji Doya, Nico Scherf
The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.
By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)
The paper discusses how reinforcement learning theory relies on probability theory via Markov chains and highlights a deep link between probability theory and potential theory. It reviews this connection and examines how a potential-theoretic perspective can be applied to core RL representations and algorithms under a fixed‑policy assumption, suggesting possible gains in sample efficiency and formal constraints. The authors also note that relaxing the fixed‑policy assumption allows the linear potential theory framework to extend naturally to nonlinear cases.
By Christopher Connolly
arXiv:2602. 20220v2 Announce Type: replace-cross Abstract: We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots.
By Yarden As, Dhruva Tirumala, Ren\'e Zurbr\"ugg, Chenhao Li, Stelian Coros, Andreas Krause, Markus Wulfmeier
We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the default reinforcement learning algorithm at OpenAI because of its ease of use and good performance.
arXiv:2608. 02433v1 Announce Type: new Abstract: Reinforcement learning and control theory are two adjacent scientific fields that focus on optimizing the controller of unknown dynamical systems using feedback.
By Claire Vernade, Onno Eberhard, Martha White, Florian D\"orfler, Csaba Szepesv\'ari, Miroslav Krstic, Michael Muehlebach