arXiv Machine Learning

Convex-Concave Reinforcement Learning

The paper introduces Convex-Concave Reinforcement Learning (CCRL), showing that the exact per‑iteration objective in policy learning can be expressed as a difference‑of‑convex (DC) program in log‑density‑ratio coordinates. This formulation unifies existing methods such as CPI, NPG, TRPO, and AWR as special cases and enables a multi‑step axis that couples consecutive decisions. Using sequential convex programming, the authors provide convergence guarantees and demonstrate that CCRL outperforms or matches PPO on diagnostic MDPs, classic control tasks, and a stochastic mid‑horizon healthcare domain, achieving faster convergence and higher training‑curve area.

arXiv Machine Learning
Jul 28

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.

By Asha Barua, Sajad Khodadadian
arXiv Machine Learning
Sep 30

Global Optimality for Constrained Exploration via Penalty Regularization

The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.

By Florian Wolf, Ilyas Fatkhullin, Niao He
arXiv Machine Learning
Sep 17

A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

The paper presents a convergence framework for deep $V$‑learning over a finite horizon $H$, deriving explicit bounds on policy loss by decomposing the Bellman update error into six residuals. It shows how $L^s$ concentrability controls expected $L^1$ loss, quantifies the impact of shared sampling across horizon levels, and provides optimal and near‑optimal sample allocations for statistical error rates. The work also establishes sharp action‑gap bounds under a margin condition, transfers optimal‑gap results to frozen‑iterate gaps, and offers consistency guarantees for generative‑reset approximate‑ERM procedures with exact action scores.

By Yury Kolomeytsev
arXiv Machine Learning
Oct 2

Fast Regularized Policy Mirror Descent with One-Step TD Updates

The paper introduces Fast Regularized Policy Mirror Descent (PMD) that pairs policy updates with a single temporal-difference (TD) critic step. It proves global linear convergence for finite discounted MDPs using exact coordinate-wise Bellman updates and any positive actor stepsize, regardless of critic initialization. For stochastic TD-PMD with strongly convex mirror maps, the authors achieve an expected value gap of ε after “~O(1/((1-γ)^5 σ_b ε))” transitions, without requiring trajectory resets, generative models, or nested policy evaluation loops.

By Qipei Chen, Wenye Li, Yule Sun, Ke Wei
arXiv Machine Learning
Aug 11

Directional-Clamp PPO

arXiv:2511. 02577v2 Announce Type: replace Abstract: Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness and effectiveness across a range of problems.

By Gilad Karpel, Ruida Zhou, Shoham Sabach, Mohammad Ghavamzadeh