arXiv Machine Learning

Rates of Convergence in the Central Limit Theorem for Markov Chains, with an Application to TD Learning

arXiv AI
Jul 17

Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.

By Mohsen Amiri, Sindri Magn\'usson