Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients
Read the original on arXiv Machine Learning →The paper investigates a natural gradient method based on the Fisher information matrix of state-action distributions, which follows a Fisher‑Rao gradient flow within the state-action polytope under a linear potential. It establishes linear convergence rates for Fisher‑Rao gradient flows of linear programs, with the rate tied to the program’s geometry, and provides improved error bounds for entropic regularization. Additionally, the authors extend their analysis to perturbed flows, proving sublinear convergence for both perturbed Fisher‑Rao and natural gradient flows, thereby encompassing state‑action natural policy gradients.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.