arXiv AI

Target Updates May Stabilize Linear Q-Learning: Periodic and Soft Dynamics

arXiv:2606. 02645v1 Announce Type: cross Abstract: Periodic target updates in Q-learning and soft target updates in actor-critic methods are empirically well established stabilization mechanisms, but their precise theoretical explanation is still incomplete.

arXiv AI
Jul 17

Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.

By Mohsen Amiri, Sindri Magn\'usson
arXiv Machine Learning
Sep 4

Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function Approximation

The paper investigates model‑free robust Q‑learning with χ² uncertainty sets and linear function approximation, using data from a single trajectory of an unknown nominal MDP. It introduces a variational reformulation of the robust Bellman target and a blockwise frozen‑target scheme to overcome estimation and non‑contractivity challenges, and proves a finite‑time error bound for every discount factor γ in (0,1). A neural‑network experiment demonstrates the practical use of the variational target in a continuous‑state nonlinear‑control task.

By Saptarshi Mandal, Yashaswini Murthy, R. Srikant