arXiv:2604. 23130v2 Announce Type: replace-cross Abstract: Jailbreak attacks expose a persistent failure mode in safety-aligned LLMs: models can be pushed into harmful behavior, but the internal representations enabling this shift remain poorly localized.
By Nilanjana Das, Mathew Dawit, Aman Chadha, Manas Gaur
arXiv:2607. 00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies.
By Shei Pern Chua, Fangzhao Wu
arXiv:2606. 03486v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translation, encoding, adversarial suffixes, and multi-turn buildup.
By Zhongyang Lin, Ziran Zhao, Feifei Zhai, Pengyuan Liu
arXiv:2606. 16242v1 Announce Type: new Abstract: The Rapid Response (RR) framework, deployed in production systems, including Anthropic's ASL-3 safeguards, continuously improves jailbreak-detection classifiers.
By David Huang, Jaewon Chang, Avidan Shah, Prateek Mittal, Chawin Sitawarin
arXiv:2606. 01738v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to LLMs by exploiting conversational dynamics such as gradual escalation and cross-turn coordination.
By Zhiqing Ma, Zhonghao Xu, Dong Yu, Chen Kang, Changliang Li, Pengyuan Liu
arXiv:2509. 03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs.
By Chuhan Zhang, Ye Zhang, Bowen Shi, Yuyou Gan, Tianyu Du, Shouling Ji, Dazhan Deng, Yingcai Wu