arXiv AI

Geometrically Averaged Hard Target Updates for Linear Q-Learning

arXiv:2606. 10835v1 Announce Type: cross Abstract: Periodic hard target updates are among the most common stabilization devices in modern deep Q-learning.

arXiv Machine Learning
Sep 4

Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function Approximation

The paper investigates model‑free robust Q‑learning with χ² uncertainty sets and linear function approximation, using data from a single trajectory of an unknown nominal MDP. It introduces a variational reformulation of the robust Bellman target and a blockwise frozen‑target scheme to overcome estimation and non‑contractivity challenges, and proves a finite‑time error bound for every discount factor γ in (0,1). A neural‑network experiment demonstrates the practical use of the variational target in a continuous‑state nonlinear‑control task.

By Saptarshi Mandal, Yashaswini Murthy, R. Srikant