arXiv Machine Learning By M. Forzo, E. Monzio Compagnoni, A. Russo, A. Pacchiano

A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise

Read the original on arXiv Machine Learning →

arXiv:2606. 18183v1 Announce Type: cross Abstract: Temporal difference (TD) learning with linear function approximation is a core method for policy evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 17

Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.

By Mohsen Amiri, Sindri Magn\'usson