Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require c...
arXiv:2602. 00781v2 Announce Type: replace Abstract: Online reinforcement learning in non-episodic, finite-horizon MDPs remains underexplored and is challenged by the need to estimate returns to a fixed terminal time.
arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.
arXiv:2609.36486v1 Announce Type: new Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M...
arXiv:2608. 07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ.
The paper introduces a new primal–dual algorithm for episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions. It achieves a rate‑optimal ×O(√K) regret and cumulative constraint violation, improving upon the previous ×O(K^{3/4}) bound and eliminating the need for Slater’s condition. The method combines adaptive FTRL, contracted value estimation, and an exponential Lyapunov function, enabling uniform concentration over the value function class and computational efficiency independent of the state‑space size.