arXiv Machine Learning By Saptarshi Mandal, Yashaswini Murthy, R. Srikant

Finite-Time Convergence of Distributionally Robust Q-Learning with Linear Function Approximation

Read the original on arXiv Machine Learning →

arXiv:2510. 01721v3 Announce Type: replace Abstract: Distributionally robust reinforcement learning (DRRL) seeks policies that perform well when the deployment transition model differs from the nominal model generating the data.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.