OpenAI Blog

Variance reduction for policy gradient with action-dependent factorized baselines

arXiv Machine Learning
Sep 11

Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients

The paper investigates a natural gradient method based on the Fisher information matrix of state-action distributions, which follows a Fisher‑Rao gradient flow within the state-action polytope under a linear potential. It establishes linear convergence rates for Fisher‑Rao gradient flows of linear programs, with the rate tied to the program’s geometry, and provides improved error bounds for entropic regularization. Additionally, the authors extend their analysis to perturbed flows, proving sublinear convergence for both perturbed Fisher‑Rao and natural gradient flows, thereby encompassing state‑action natural policy gradients.

By Johannes M\"uller, Semih \c{C}ayc{\i}, Guido Mont\'ufar
arXiv Machine Learning
1d ago

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

AdaStep introduces an adaptive step-credit weighting technique for agentic reinforcement learning, addressing the coarse granularity of trajectory-level objectives in long-horizon LLM agents. By formulating the weighting as a mean-squared-error estimation problem and deriving an optimal per-state shrinkage coefficient, AdaStep selectively preserves local credit when return variation is due to the chosen action and suppresses it when downstream randomness dominates. The method requires only lightweight scalar computations, no critic or extra rollouts, and demonstrates consistent performance gains across three model backbones on ALFWorld, WebShop, and ScienceWorld.

By Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas