arXiv Machine Learning By Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum

Convex-Concave Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper introduces Convex-Concave Reinforcement Learning (CCRL), showing that the exact per‑iteration objective in policy learning can be expressed as a difference‑of‑convex (DC) program in log‑density‑ratio coordinates. This formulation unifies existing methods such as CPI, NPG, TRPO, and AWR as special cases and enables a multi‑step axis that couples consecutive decisions. Using sequential convex programming, the authors provide convergence guarantees and demonstrate that CCRL outperforms or matches PPO on diagnostic MDPs, classic control tasks, and a stochastic mid‑horizon healthcare domain, achieving faster convergence and higher training‑curve area.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 28

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.

By Asha Barua, Sajad Khodadadian
arXiv Machine Learning
Sep 30

Global Optimality for Constrained Exploration via Penalty Regularization

The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.

By Florian Wolf, Ilyas Fatkhullin, Niao He