arXiv Machine Learning

Robust Asynchronous Q-Learning under Reward and State Corruption via Batching

arXiv:2607. 20822v1 Announce Type: new Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback.

arXiv Machine Learning
Sep 18

Robust Federated Q-Learning with Almost No Communication

The paper introduces Robust Fed-Q, a federated Q‑learning algorithm designed for settings where multiple agents interact with a shared Markov Decision Process and communicate through a central server. It combines model‑based and model‑free reinforcement learning techniques with a median‑of‑means strategy from robust statistics to handle a small fraction of adversarial agents. The authors prove that Robust Fed-Q achieves exact convergence to the optimal value function with high probability, attains near‑optimal finite‑time rates that benefit from collaboration, and requires only “~O(1)” communication rounds per guarantee.

By Sreejeet Maity, Aritra Mitra
arXiv Machine Learning
Sep 14

Inverting Self-Triggered Control: Adversarial Reinforcement Learning for Sparse Denial-of-Service Attacks

The paper introduces an adversarial reinforcement learning framework that learns the sparsest Denial-of-Service (DoS) attack schedule capable of destabilizing self‑triggered reinforcement learning controllers (RL‑STC). It proves a lower bound on the minimum number of jamming actions needed to force a crash and demonstrates that the learned adversary consistently defeats four different defenders—one LQR and three RL‑STC—across Pendulum, CartPole, and Quadrotor2D environments, outperforming greedy and periodic baselines in jam‑time‑per‑failure. The study also shows that the adversary remains effective under Gaussian observation noise and limited state information.

By Adam Haroon, Erick J. Rodr\'iguez-Seda, Tristan Schuler, Cody Fleming