arXiv:2607. 04645v1 Announce Type: cross Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests.
By Samira Hajizadeh
arXiv:2608. 02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Xu Peng, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
arXiv:2605. 00994v2 Announce Type: replace-cross Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors.
By Mohammed Abu Baker, Luca Baroni, Dan Wilhelm
arXiv:2607. 26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools.
By Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao, Zhixuan Chu, Wanyu Lin, Tianhang Zheng
arXiv:2608. 13250v1 Announce Type: cross Abstract: Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge.
By Long Hoang Nguyen, Brice Valentin Kok-Shun, Guangyu Du, Ali Sunyaev
arXiv:2606. 10740v1 Announce Type: new Abstract: Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation.
By Sai Kartheek Reddy Kasu, Nils Lukas, Samuele Poppi