arXiv Machine Learning

Finite-Time Convergence of Distributionally Robust Q-Learning with Linear Function Approximation

arXiv:2510. 01721v3 Announce Type: replace Abstract: Distributionally robust reinforcement learning (DRRL) seeks policies that perform well when the deployment transition model differs from the nominal model generating the data.

arXiv Machine Learning
Sep 4

Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function Approximation

The paper investigates model‑free robust Q‑learning with χ² uncertainty sets and linear function approximation, using data from a single trajectory of an unknown nominal MDP. It introduces a variational reformulation of the robust Bellman target and a blockwise frozen‑target scheme to overcome estimation and non‑contractivity challenges, and proves a finite‑time error bound for every discount factor γ in (0,1). A neural‑network experiment demonstrates the practical use of the variational target in a continuous‑state nonlinear‑control task.

By Saptarshi Mandal, Yashaswini Murthy, R. Srikant
arXiv Machine Learning
Aug 24

Smart Exploration in Reinforcement Learning using Bounded Uncertainty Models

The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.

By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes
arXiv Machine Learning
Sep 18

Robust Federated Q-Learning with Almost No Communication

The paper introduces Robust Fed-Q, a federated Q‑learning algorithm designed for settings where multiple agents interact with a shared Markov Decision Process and communicate through a central server. It combines model‑based and model‑free reinforcement learning techniques with a median‑of‑means strategy from robust statistics to handle a small fraction of adversarial agents. The authors prove that Robust Fed-Q achieves exact convergence to the optimal value function with high probability, attains near‑optimal finite‑time rates that benefit from collaboration, and requires only “~O(1)” communication rounds per guarantee.

By Sreejeet Maity, Aritra Mitra