arXiv Machine Learning

Finite-Time Convergence of Distributionally Robust Q-Learning with Linear Function Approximation

arXiv:2510. 01721v3 Announce Type: replace Abstract: Distributionally robust reinforcement learning (DRRL) seeks policies that perform well when the deployment transition model differs from the nominal model generating the data.