arXiv Machine Learning

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation

arXiv:2605. 18591v2 Announce Type: replace Abstract: Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the Fisher matrix.

arXiv Machine Learning
Sep 11

Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients

The paper investigates a natural gradient method based on the Fisher information matrix of state-action distributions, which follows a Fisher‑Rao gradient flow within the state-action polytope under a linear potential. It establishes linear convergence rates for Fisher‑Rao gradient flows of linear programs, with the rate tied to the program’s geometry, and provides improved error bounds for entropic regularization. Additionally, the authors extend their analysis to perturbed flows, proving sublinear convergence for both perturbed Fisher‑Rao and natural gradient flows, thereby encompassing state‑action natural policy gradients.

By Johannes M\"uller, Semih \c{C}ayc{\i}, Guido Mont\'ufar
arXiv AI
Sep 18

Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design

The paper investigates how reinforcement learning can be effectively applied to diffusion models for visual tasks, focusing on the role of likelihood estimation. By systematically separating policy‑gradient objectives, likelihood estimators, and rollout sampling schemes, the authors find that using an evidence lower bound (ELBO) based likelihood estimator computed from the final generated sample is the key factor for stable and efficient RL optimization, outweighing the choice of loss function. Experiments on SD 3.5 Medium across multiple reward benchmarks confirm that this approach improves GenEval scores from 0.24 to 0.95 in 90 GPU hours, outperforming existing methods such as FlowGRPO and the current state‑of‑the‑art without reward hacking.

By Jaemoo Choi, Yuchen Zhu, Wei Guo, Petr Molodyk, Bo Yuan, Jinbin Bai, Yi Xin, Molei Tao, Yongxin Chen
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas
arXiv Machine Learning
Jun 25

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

arXiv:2601. 23075v2 Announce Type: replace Abstract: On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy, and policy updates must be conservative.

By Yuexin Bian, Jie Feng, Tao Wang, Yijiang Li, Sicun Gao, Yuanyuan Shi
arXiv Machine Learning
Aug 5

Information-Geometric Forward Policy Training in GFlowNets

arXiv:2608. 03967v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward.

By Yordan Raykov, Rodrigo Veiga
arXiv Machine Learning
Jun 5

On Advantage Estimates for Max@K Policy Gradients

arXiv:2606. 06080v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards is widely used for post-training reasoning models, but sparse outcome rewards make exploration difficult.

By Shota Takashiro, Soichiro Nishimori, Paavo Parmas, Yongmin Kim, Kohsei Matsutani, Gouki Minegishi, Yusuke Iwasawa, Takeshi Kojima, Yutaka Matsuo