In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solvi...
arXiv:2512. 14617v2 Announce Type: replace-cross Abstract: Many practical decision-making problems involve tasks whose success depends on the entire system history, rather than on achieving a state with desired properties.
By Alessandro Trapasso, Luca Iocchi, Fabio Patrizi
The paper introduces a new approach to learning chance-constrained Markov decision processes (CCMDPs) using a Bellman distributional certificate. It provides both model-based and model-free algorithms with theoretical guarantees, including matching upper and lower bounds for tabular discounted CCMDPs with bounded successor support. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage benchmark demonstrate the safety and effectiveness of the proposed methods.
By Chenbei Lu, Hongyu Yi
arXiv:2606. 10979v1 Announce Type: new Abstract: Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operational constraints.
By Yi Chen (Lucy), Rushuai Yang (Lucy), Qiang Chen (Lucy), Dongyan (Lucy), Huo
The paper introduces reinforcement learning for Continuous-Time Jump Markov Decision Processes (CTJMDPs) with general discrete state spaces and continuous/discrete actions. It develops entropy‑regularized continuous‑time control and establishes theoretical foundations for q‑learning in this setting, providing model‑free algorithms that outperform naive discretization. Numerical tests on network dynamic pricing demonstrate the method’s ability to learn near‑optimal policies and scale to large networks.
By Huiling Meng, Ningyuan Chen, Xuefeng Gao
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.
By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.
By Florian Wolf, Ilyas Fatkhullin, Niao He
The paper introduces a computationally efficient algorithm for infinite-horizon average-reward constrained Markov decision processes (CMDPs) under weak communication. It achieves a high-probability regret and cumulative constraint violation of ×O(√T) in the tabular setting, matching optimal dependence up to logarithmic factors. The method augments the state with cumulative constraint violation, reshapes rewards using a Huber potential, and applies finite-horizon approximation with optimistic value iteration to maintain bounded per-step rewards.
By Kihyun Yu, Seoungbin Bae, Dabeen Lee
The paper introduces Convex-Concave Reinforcement Learning (CCRL), showing that the exact per‑iteration objective in policy learning can be expressed as a difference‑of‑convex (DC) program in log‑density‑ratio coordinates. This formulation unifies existing methods such as CPI, NPG, TRPO, and AWR as special cases and enables a multi‑step axis that couples consecutive decisions. Using sequential convex programming, the authors provide convergence guarantees and demonstrate that CCRL outperforms or matches PPO on diagnostic MDPs, classic control tasks, and a stochastic mid‑horizon healthcare domain, achieving faster convergence and higher training‑curve area.
By Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum
arXiv:2602. 00781v2 Announce Type: replace Abstract: Online reinforcement learning in non-episodic, finite-horizon MDPs remains underexplored and is challenged by the need to estimate returns to a fixed terminal time.
By Jiamin Xu, Kyra Gan
arXiv:2609.24489v1 Announce Type: new
Abstract: Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target n...
By Hyukjun Yang, Jongchan Park, Narim Jeong, Donghwan Lee