arXiv Machine Learning

Inverting Self-Triggered Control: Adversarial Reinforcement Learning for Sparse Denial-of-Service Attacks

The paper introduces an adversarial reinforcement learning framework that learns the sparsest Denial-of-Service (DoS) attack schedule capable of destabilizing self‑triggered reinforcement learning controllers (RL‑STC). It proves a lower bound on the minimum number of jamming actions needed to force a crash and demonstrates that the learned adversary consistently defeats four different defenders—one LQR and three RL‑STC—across Pendulum, CartPole, and Quadrotor2D environments, outperforming greedy and periodic baselines in jam‑time‑per‑failure. The study also shows that the adversary remains effective under Gaussian observation noise and limited state information.

arXiv AI
Sep 21

Taming the Adversary: A Cost-to-Disturbance Ratio Approach to Adversarial Reinforcement Learning

The paper introduces CoDRA, a cost-to-disturbance ratio approach for adversarial reinforcement learning that balances controller performance and disturbance exposure without extra penalty terms. CoDRA uses a self‑normalized actor–critic update, scaling value terms by a stop‑gradient normalization constant derived from the current batch. Experiments on MuJoCo pendulum tasks show that CoDRA achieves the lowest cost across a range of forces and masses, outperforming other methods especially on the more challenging InvertedDoublePendulum environment.

By Taeho Lee, Donghwan Lee
arXiv AI
Jun 18

TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning

arXiv:2606. 18308v1 Announce Type: cross Abstract: Safe coordination in networked cyber-physical systems forces learning algorithms to simultaneously handle hybrid discrete-continuous actions, hard training-time safety constraints, and physics-governed dynamics.

By Zijie Meng, Ziwei Li, Yufei Liu, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Miao Zhang
arXiv Machine Learning
Aug 17

Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty

arXiv:2608. 13651v1 Announce Type: cross Abstract: We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-\Delta, \Delta]$ is unknown in sign and magnitude and $\Delta$ is arbitrarily large.

By Dimitar Ho
arXiv AI
Jun 18

A Distributionally Robust Reinforcement Learning Framework for Constrained Urban EV Dispatch

arXiv:2604. 25848v2 Announce Type: replace Abstract: We study city-scale control of electric-vehicle (EV) ride-hailing fleets where dispatch, repositioning, and charging decisions must respect charger and feeder limits under uncertain, spatially correlated demand and travel times.

By An Nguyen, Hoang Nguyen, Phuong Le, Hung Pham, Cuong Do, Laurent El Ghaoui
arXiv Machine Learning
1d ago

Rate-Optimal Algorithm for Adversarial Linear CMDPs

The paper introduces a new primal–dual algorithm for episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions. It achieves a rate‑optimal ×O(√K) regret and cumulative constraint violation, improving upon the previous ×O(K^{3/4}) bound and eliminating the need for Slater’s condition. The method combines adaptive FTRL, contracted value estimation, and an exponential Lyapunov function, enabling uniform concentration over the value function class and computational efficiency independent of the state‑space size.

By Kihyun Yu, Honghao Wei, Dabeen Lee
arXiv Machine Learning
4d ago

Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

The paper introduces a curriculum reinforcement learning approach to overcome the cold‑start problem in prompt‑injection red‑teaming of frontier large language models. By training an attacker LLM sequentially against increasingly robust target models and ensuring partial success at each stage, the method achieves high attack success rates (93.8% against GPT‑5.6‑Luna and 45.0% against GPT‑5.6‑Terra) where prior RL methods fail. The attacker LLM also transfers its effectiveness to other frontier models it was not explicitly trained on.

By Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia