Variance reduction for policy gradient with action-dependent factorized baselines
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
arXiv:2602. 05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization.
The paper investigates a natural gradient method based on the Fisher information matrix of state-action distributions, which follows a Fisher‑Rao gradient flow within the state-action polytope under a linear potential. It establishes linear convergence rates for Fisher‑Rao gradient flows of linear programs, with the rate tied to the program’s geometry, and provides improved error bounds for entropic regularization. Additionally, the authors extend their analysis to perturbed flows, proving sublinear convergence for both perturbed Fisher‑Rao and natural gradient flows, thereby encompassing state‑action natural policy gradients.
arXiv:2609.06882v1 Announce Type: cross Abstract: Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remain...
arXiv:2511. 23310v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective paradigm for post-training large language models, yet the design of its baselines and learning-rate schedules remains largely heuristic.
arXiv:2609.14327v1 Announce Type: new Abstract: Variance penalization is a principled approach to risk-sensitive reinforcement learning (RL) that explicitly trades expected return for policy stabilit...
AdaStep introduces an adaptive step-credit weighting technique for agentic reinforcement learning, addressing the coarse granularity of trajectory-level objectives in long-horizon LLM agents. By formulating the weighting as a mean-squared-error estimation problem and deriving an optimal per-state shrinkage coefficient, AdaStep selectively preserves local credit when return variation is due to the chosen action and suppresses it when downstream randomness dominates. The method requires only lightweight scalar computations, no critic or extra rollouts, and demonstrates consistent performance gains across three model backbones on ALFWorld, WebShop, and ScienceWorld.