arXiv:2606. 10580v1 Announce Type: cross Abstract: The asymptotic behaviour of Monte Carlo optimistic policy iteration (MC-O-PI) is a long-standing open question.
By Octave Oliviers, Glenn Vinnicombe
arXiv:2610.00911v1 Announce Type: new
Abstract: We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future...
By Taehyun Hwang, Hyunjun Choi, Heesang Ann, Min-hwan Oh
arXiv:2609.39837v1 Announce Type: new
Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
By Qipei Chen, Wenye Li, Yule Sun, Ke Wei
arXiv:2609.39093v1 Announce Type: new
Abstract: We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probabi...
By Kihyun Yu, Seoungbin Bae, Dabeen Lee
arXiv:2606. 15247v1 Announce Type: cross Abstract: The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting.
By Octave Oliviers, Glenn Vinnicombe
The paper investigates the finite‑iteration behavior of exact asynchronous recursions used in categorical distributional temporal‑difference (TD) learning. It analyzes both scalar categorical TD in the Cramér geometry and multivariate signed‑categorical TD in the maximum mean discrepancy geometry, showing that these methods can be viewed as single‑state stochastic‑approximation recursions that contract in a block‑supremum norm. The authors develop a restricted‑domain theory, derive discounted bounds under i.i.d. and Markovian sampling, and extend the analysis to undiscounted fixed‑horizon policy evaluation with horizon‑stacked categorical methods under episodic sampling, thereby providing a unified non‑asymptotic analysis across various settings.
By Ege C. Kaya, Abolfazl Hashemi