The paper introduces Robust Fed-Q, a federated Q‑learning algorithm designed for settings where multiple agents interact with a shared Markov Decision Process and communicate through a central server. It combines model‑based and model‑free reinforcement learning techniques with a median‑of‑means strategy from robust statistics to handle a small fraction of adversarial agents. The authors prove that Robust Fed-Q achieves exact convergence to the optimal value function with high probability, attains near‑optimal finite‑time rates that benefit from collaboration, and requires only “~O(1)” communication rounds per guarantee.
By Sreejeet Maity, Aritra Mitra
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).
By Yang Xu, Swetha Ganesh, Vaneet Aggarwal
arXiv:2605. 28276v2 Announce Type: replace Abstract: Reinforcement learning algorithms are commonly analyzed (and designed) under the Markov assumption.
By Onno Eberhard, Claire Vernade, Michael Muehlebach
arXiv:2605. 01752v4 Announce Type: replace Abstract: We study linear dueling bandits in volatile environments characterized by the simultaneous presence of post-serving contexts, delayed feedback, and adversarial corruption.
By Youngmin Oh
arXiv:2402. 06734v2 Announce Type: replace-cross Abstract: We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting.
By Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban, Adish Singla, Goran Radanovi\'c
arXiv:2606. 03521v1 Announce Type: cross Abstract: To improve the real-world applicability of reinforcement learning (RL), the field of adversarially robust RL studies how to train agents under adversarial environment perturbations.
By Siemen Herremans, Ali Anwar, Siegfried Mercelis