arXiv Machine Learning By Vasos Arnaoutis, Eric Lutters, Bojana Rosi\'c

Generalized Kalman filter based temporal difference reinforcement learning

Read the original on arXiv Machine Learning →

arXiv:2607. 20010v1 Announce Type: new Abstract: In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 17

Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.

By Mohsen Amiri, Sindri Magn\'usson