arXiv:2607. 28390v1 Announce Type: new Abstract: Constrained Markov Decision Processes (CMDPs) provide a natural framework for reinforcement learning in safety-critical applications, where agents maximize long-term reward while satisfying long-term constraints.
By Ankur Naskar, Vaneet Aggarwal
The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.
By Florian Wolf, Ilyas Fatkhullin, Niao He
The paper introduces a new primal–dual algorithm for episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions. It achieves a rate‑optimal ×O(√K) regret and cumulative constraint violation, improving upon the previous ×O(K^{3/4}) bound and eliminating the need for Slater’s condition. The method combines adaptive FTRL, contracted value estimation, and an exponential Lyapunov function, enabling uniform concentration over the value function class and computational efficiency independent of the state‑space size.
By Kihyun Yu, Honghao Wei, Dabeen Lee
arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.
By Umer Siddique, Peilang Li, Yongcan Cao
arXiv:2505. 15201v5 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) algorithms sample multiple n>1 solution attempts for each problem and reward them independently.
By Christian Walder, Deep Karkhanis
arXiv:2605. 11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories.
By Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz
The paper introduces Exchange Policy Optimization (EPO), a framework for semi‑infinite safe reinforcement learning that handles infinitely many constraints by iteratively solving finite subproblems. EPO expands or deletes constraints based on tolerance violations and Lagrange multipliers, maintaining computational tractability while converging to an optimal policy with bounded safety violations. The authors prove finite convergence, provide iteration bounds, and quantify the suboptimality gap under mild assumptions.
By Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li
arXiv:2608. 02343v1 Announce Type: cross Abstract: Many operational problems are constrained sequential decision processes with large, combinatorial action spaces and interdependent feasibility constraints.
By Patrick Helm, Jan-Niklas Doerr, Joren Gijsbrechts, Stefan Minner
BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.
By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.
By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz
arXiv:2608. 10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints.
By Chenhua Fan, Jiahui Zhu, Yuhang Zhang, Honghao Wei
arXiv:2606. 08779v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or even training collapses.
By Jiashun Liu, Runze Liu, Xu Wan, Jing Liang, Hongyao Tang, Ling Pan