The paper introduces CoDRA, a cost-to-disturbance ratio approach for adversarial reinforcement learning that balances controller performance and disturbance exposure without extra penalty terms. CoDRA uses a self‑normalized actor–critic update, scaling value terms by a stop‑gradient normalization constant derived from the current batch. Experiments on MuJoCo pendulum tasks show that CoDRA achieves the lowest cost across a range of forces and masses, outperforming other methods especially on the more challenging InvertedDoublePendulum environment.
By Taeho Lee, Donghwan Lee
arXiv:2607. 06643v1 Announce Type: cross Abstract: Backdoor attacks severely threaten large-scale AI models.
By Issam Seddik, Sami Souihi, Mohamed Tamaazousti, Sara Tucci Piergiovanni
arXiv:2607. 20822v1 Announce Type: new Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback.
By Sreejeet Maity, Aritra Mitra
arXiv:2607. 27626v1 Announce Type: new Abstract: Safety-critical IoT systems such as industrial closed-loop control, V2X coordination, and remote teleoperation require every sensor's peak Age of Information (peak AoI, also abbreviated PAoI) to stay below a hard per-slot deadline, not merely an average bound.
By Wentao Zhang, Wentao Mo
arXiv:2607. 16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself.
By Venkatesh Saligrama
arXiv:2606. 18308v1 Announce Type: cross Abstract: Safe coordination in networked cyber-physical systems forces learning algorithms to simultaneously handle hybrid discrete-continuous actions, hard training-time safety constraints, and physics-governed dynamics.
By Zijie Meng, Ziwei Li, Yufei Liu, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Miao Zhang
arXiv:2608. 06520v1 Announce Type: new Abstract: We study online cooperative control of a multi-agent system under Byzantine attacks.
By Ximing Sun, Yue Wang
arXiv:2608. 13651v1 Announce Type: cross Abstract: We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-\Delta, \Delta]$ is unknown in sign and magnitude and $\Delta$ is arbitrarily large.
By Dimitar Ho
arXiv:2601. 07674v2 Announce Type: replace-cross Abstract: Random walk (RW)-based algorithms have long been popular in distributed systems due to low overheads and scalability, with recent growing applications in decentralized learning.
By Xingran Chen, Parimal Parag, Rohit Bhagat, Salim El Rouayheb
arXiv:2604. 25848v2 Announce Type: replace Abstract: We study city-scale control of electric-vehicle (EV) ride-hailing fleets where dispatch, repositioning, and charging decisions must respect charger and feeder limits under uncertain, spatially correlated demand and travel times.
By An Nguyen, Hoang Nguyen, Phuong Le, Hung Pham, Cuong Do, Laurent El Ghaoui
The paper introduces a new primal–dual algorithm for episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions. It achieves a rate‑optimal ×O(√K) regret and cumulative constraint violation, improving upon the previous ×O(K^{3/4}) bound and eliminating the need for Slater’s condition. The method combines adaptive FTRL, contracted value estimation, and an exponential Lyapunov function, enabling uniform concentration over the value function class and computational efficiency independent of the state‑space size.
By Kihyun Yu, Honghao Wei, Dabeen Lee
The paper introduces a curriculum reinforcement learning approach to overcome the cold‑start problem in prompt‑injection red‑teaming of frontier large language models. By training an attacker LLM sequentially against increasingly robust target models and ensuring partial success at each stage, the method achieves high attack success rates (93.8% against GPT‑5.6‑Luna and 45.0% against GPT‑5.6‑Terra) where prior RL methods fail. The attacker LLM also transfers its effectiveness to other frontier models it was not explicitly trained on.
By Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia