arXiv:2606. 25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints.
By Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal
arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.
By Asha Barua, Sajad Khodadadian
The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.
By Florian Wolf, Ilyas Fatkhullin, Niao He
arXiv:2512. 23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$.
By Yingru Li, Jiacai Liu, Jiawei Xu, Yuxuan Tong, Ziniu Li, Qian Liu, Baoxiang Wang
arXiv:2610.09577v1 Announce Type: new
Abstract: We study finite-horizon online resource allocation with i.i.d. requests and an endogenous Markov state on a finite state space: each action affects the...
By Zhaohua Chen
arXiv:2609.36486v1 Announce Type: new
Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M...
By Zijun Chen, Zihan Zhang
arXiv:2609.37660v1 Announce Type: new
Abstract: We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each roun...
By Wansoo Choi, Seoungbin Bae, Dabeen Lee
The paper presents a convergence framework for deep $V$‑learning over a finite horizon $H$, deriving explicit bounds on policy loss by decomposing the Bellman update error into six residuals. It shows how $L^s$ concentrability controls expected $L^1$ loss, quantifies the impact of shared sampling across horizon levels, and provides optimal and near‑optimal sample allocations for statistical error rates. The work also establishes sharp action‑gap bounds under a margin condition, transfers optimal‑gap results to frozen‑iterate gaps, and offers consistency guarantees for generative‑reset approximate‑ERM procedures with exact action scores.
By Yury Kolomeytsev
The paper introduces Fast Regularized Policy Mirror Descent (PMD) that pairs policy updates with a single temporal-difference (TD) critic step. It proves global linear convergence for finite discounted MDPs using exact coordinate-wise Bellman updates and any positive actor stepsize, regardless of critic initialization. For stochastic TD-PMD with strongly convex mirror maps, the authors achieve an expected value gap of ε after “~O(1/((1-γ)^5 σ_b ε))” transitions, without requiring trajectory resets, generative models, or nested policy evaluation loops.
By Qipei Chen, Wenye Li, Yule Sun, Ke Wei
arXiv:2511. 02577v2 Announce Type: replace Abstract: Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness and effectiveness across a range of problems.
By Gilad Karpel, Ruida Zhou, Shoham Sabach, Mohammad Ghavamzadeh
arXiv:2610.08384v1 Announce Type: new
Abstract: In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT con...
By Zihao Zhao, Ashwath K. Karunakaram, Ali Eshragh, Yuexing Li, Kai Wang
arXiv:2604. 06039v2 Announce Type: replace-cross Abstract: Value iteration-type methods have been extensively studied for computing a nearly optimal value function in reinforcement learning (RL).
By Zhichao Jia, Guanghui Lan