arXiv Machine Learning

Model-Adaptive and Risk-Constrained Frequency Hopping Against Predictive Jammers

arXiv Machine Learning
4d ago

Restless Bandits with Individual Penalty Constraints: Near-Optimal Indices and Deep Reinforcement Learning

This paper studies Restless Multi‑Armed Bandits with individual penalty constraints for dynamic wireless networks, allowing each arm to have distinct performance limits such as energy, activation, or age of information. It introduces the Penalty‑Optimal Whittle (POW) index, which depends only on an arm’s transition kernel and its constraints, making it computable offline and independent of system‑wide parameters. The authors prove the POW index policy is asymptotically optimal, present a deep reinforcement learning method to learn the index online, and show through simulations that it outperforms existing policies.

By Nida Zamir, I-Hong Hou
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv Machine Learning
1d ago

Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift

The paper studies how an eavesdropper can adaptively attack quantum key distribution (QKD) systems when channel noise and device drift vary over time. By modeling the attack as a constrained Markov decision process and using reinforcement learning to jointly search gate structures and rotation angles, the authors construct compact attack circuits that perform near the theoretical upper bound for both device‑independent E91 and BB84 protocols under realistic noise models. The results show that adaptive attacks can significantly increase the eavesdropper’s information compared to fixed‑circuit strategies, and that the learned attacks recover known optimal cloners and key‑rate bounds.

By Marcel Mordarski, Benjamin Gras, Abdelrahman Shehata, Daniel Budina, Roberto Bondesan
arXiv Machine Learning
Sep 14

Inverting Self-Triggered Control: Adversarial Reinforcement Learning for Sparse Denial-of-Service Attacks

The paper introduces an adversarial reinforcement learning framework that learns the sparsest Denial-of-Service (DoS) attack schedule capable of destabilizing self‑triggered reinforcement learning controllers (RL‑STC). It proves a lower bound on the minimum number of jamming actions needed to force a crash and demonstrates that the learned adversary consistently defeats four different defenders—one LQR and three RL‑STC—across Pendulum, CartPole, and Quadrotor2D environments, outperforming greedy and periodic baselines in jam‑time‑per‑failure. The study also shows that the adversary remains effective under Gaussian observation noise and limited state information.

By Adam Haroon, Erick J. Rodr\'iguez-Seda, Tristan Schuler, Cody Fleming