arXiv:2609.26185v1 Announce Type: cross
Abstract: Large Language Models (LLMs) demonstrate impressive capabilities across many applications but remain vulnerable to jailbreak attacks, which elicit ha...
By Doniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen Yang
The paper introduces a self‑evolving defense framework for large language models that uses a persistent, cross‑interaction rule memory to adapt to new jailbreak attacks. When an attack succeeds, the system abstracts the failure into a method‑level rule that captures the structural attack wrapper, allowing the rule to generalize across an entire attack family. This memory‑based adaptation operates without parameter updates, works with both open‑weight and black‑box models, and has been shown to reduce attack success rates while preserving benign utility across multiple jailbreak families.
By Tongyan Hu, Bryan Hooi
The paper "Jailbreaking in the Haystack" introduces NINJA, a jailbreak technique that exploits long-context language models by appending benign, model-generated content to harmful user goals. It demonstrates that the position of harmful goals within the context is crucial for safety, and shows that NINJA significantly boosts attack success rates on models such as LLaMA, Qwen, Mistral, and Gemini. Unlike previous methods, NINJA is low-resource, transferable, less detectable, and compute‑optimal, revealing that carefully crafted benign long contexts can expose fundamental vulnerabilities in modern LMs.
By Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan
arXiv:2508. 10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks.
By Jinhwa Kim, Ian G. Harris
arXiv:2502. 09755v4 Announce Type: replace-cross Abstract: Safety-aligned LLMs respond to prompts with either compliance or refusal, each corresponding to distinct directions in the model's activation space.
By Amit Levi, Rom Himelstein, Yaniv Nemcovsky, Avi Mendelson, Chaim Baskin
arXiv:2606. 11425v1 Announce Type: cross Abstract: Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can adapt but often relies on low-level mutations that require many target queries.
By Ge Shi, Jun Yin, Donglin Xie, Fangyi Liu, Yucan Li, Menglin Liu