arXiv Machine Learning
Oct 2

Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation

The paper introduces a computationally efficient algorithm for infinite-horizon average-reward constrained Markov decision processes (CMDPs) under weak communication. It achieves a high-probability regret and cumulative constraint violation of ×O(√T) in the tabular setting, matching optimal dependence up to logarithmic factors. The method augments the state with cumulative constraint violation, reshapes rewards using a Huber potential, and applies finite-horizon approximation with optimistic value iteration to maintain bounded per-step rewards.

By Kihyun Yu, Seoungbin Bae, Dabeen Lee
arXiv Machine Learning
Sep 24

Limiting-Kernel Q($\lambda$): Bridging Short and Long Horizons

Limiting‑Kernel Q(λ) (LKQL) is an off‑policy value estimator that blends n‑step truncation with a long‑horizon approximation based on the limiting kernel. It maintains the computational efficiency of n‑step methods while improving policy evaluation accuracy, especially for long‑horizon tasks. The authors prove faster convergence of LKQL’s operator under aperiodicity and near‑on‑policy conditions, and demonstrate empirical gains on MuJoCo continuous‑control benchmarks.

By Tolga Ok, Arman Sharifi Kolarijani, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi Kolarijani