The paper introduces Robust Fed-Q, a federated Q‑learning algorithm designed for settings where multiple agents interact with a shared Markov Decision Process and communicate through a central server. It combines model‑based and model‑free reinforcement learning techniques with a median‑of‑means strategy from robust statistics to handle a small fraction of adversarial agents. The authors prove that Robust Fed-Q achieves exact convergence to the optimal value function with high probability, attains near‑optimal finite‑time rates that benefit from collaboration, and requires only “~O(1)” communication rounds per guarantee.
By Sreejeet Maity, Aritra Mitra
arXiv:2607. 17823v1 Announce Type: new Abstract: Reinforcement Learning is a cornerstone technique for modern large reasoning models.
By Riccardo Poiani, Martino Bernasconi, Andrea Celli
arXiv:2607. 20822v1 Announce Type: new Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback.
By Sreejeet Maity, Aritra Mitra
arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.
By Volodymyr Tkachuk, Csaba Szepesv\'ari, Xiaoqi Tan
The paper introduces state abstractions that preserve the difference of Q‑functions for offline reinforcement learning, aiming to exclude irrelevant dynamics from rich state data. It proposes a dynamic generalization of the R‑learner that uses orthogonal estimation and sparse learning to estimate the Q‑function contrast, achieving faster convergence and consistency under a margin condition. Experiments on simulated and simulator‑augmented real data show variance reductions and demonstrate that the necessary information for sequential decision‑making can be smaller than that required for full state prediction.
By Defu Cao, Angela Zhou
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes