arXiv:2610.00861v1 Announce Type: cross
Abstract: Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget s...
By Melika Shirian, Kianoosh Vadaei
arXiv:2605.23220v2 Announce Type: replace
Abstract: Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated...
By Zhixiang Guo, Siyuan Liang, Shi Fu, Cheng Guo, Andras Balogh, Mark Jelasity, Dacheng Tao
arXiv:2606. 16605v1 Announce Type: new Abstract: World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making.
By Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu
arXiv:2604. 26360v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) systems face a compounding alignment challenge: not only are learned reward models uncertain about unseen state-action pairs, but the human preference annotations they are trained on are themselves inconsistent, context-dependent, and noisy.
By Disha Singha
arXiv:2607. 10630v1 Announce Type: cross Abstract: Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data.
By Tong Nie, Yuewen Mei, Junlin He, Yihong Tang, Jian Sun, Wei Ma
arXiv:2609.38938v1 Announce Type: new
Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies...
By Xinyi Ni, Lifeng Lai
arXiv:2608. 00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another agent can search for responses.
By Alina Kapanova, Arun Kanhai, Natan Vidra, Spurthi Setty
arXiv:2602. 05459v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method.
By Jan Malte T\"opperwien, Aditya Mohan, Marius Lindauer
The paper investigates a vulnerability in feedback‑based agent planning, showing that the first round of feedback corrects a large portion of adversarial directions (46%) while subsequent rounds see a sharp decline (13% and 7%). The authors attribute this to an initialization anchoring weakness driven by plausible plan shifts, lack of counterevidence, and persistence of accepted directions. They introduce “InitAnchor”, a black‑box attack framework that exploits these factors, achieving high attack success rates across diverse tasks, architectures, and LLMs, and remaining effective against multiple defenses and real‑world agents.
By Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang, Xinyu Gao, Peng Zhan, Zheng Li, Shanqing Guo
arXiv:2603. 25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs).
By John Ayotunde, Qinghua Xu, Guancheng Wang, Lionel C. Briand
arXiv:2607. 09993v1 Announce Type: cross Abstract: Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs, require strategies that are robust to hidden opponent types, such as a hidden goal flag, and to deception.
By Naman Aggarwal, Jonathan P. How
The paper investigates how an informed adversary can influence the optimal signal in a constrained signalling channel. It finds that the adversary‑robust optimum aligns with the salience pole on most items, differing only on a small subset where the salience‑to‑Bayes coordinate is undefined. The study demonstrates that as the adversary’s persuasion budget increases, the optimal signal shifts from a posterior‑maximizing to a margin‑maximizing strategy, and provides a diagnostic check for evaluating adversary‑awareness.
By Cris Huynh