arXiv:2609.39902v1 Announce Type: cross
Abstract: Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attack...
By Zhen Liang, Hai Huang, Wentao Chen
arXiv:2503. 24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers.
By Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, Ruiyuan Xu, Zhicheng Li, Yangyu Zhang, Shuaijiang Li, Yuan Wen, Chunwei Xia, Zheng Wang, Xiaobing Feng, Huimin Cui
The paper "Jailbreaking in the Haystack" introduces NINJA, a jailbreak technique that exploits long-context language models by appending benign, model-generated content to harmful user goals. It demonstrates that the position of harmful goals within the context is crucial for safety, and shows that NINJA significantly boosts attack success rates on models such as LLaMA, Qwen, Mistral, and Gemini. Unlike previous methods, NINJA is low-resource, transferable, less detectable, and compute‑optimal, revealing that carefully crafted benign long contexts can expose fundamental vulnerabilities in modern LMs.
By Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan
arXiv:2510.17904v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) are widely used because they process structures, syntax and code well, but this same ability also makes them par...
By Amirkia Rafiei Oskooei, Mehmet S. Aktas
arXiv:2607. 00481v1 Announce Type: cross Abstract: Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs).
By Junlong Liu, Haobo Wang, Weiqi Luo, Xiaojun Jia
arXiv:2606. 28962v1 Announce Type: cross Abstract: Model quantization is essential for the efficient deployment of Large Language Models (LLMs), but introduces a critical vulnerability: Quantization-Conditioned Backdoor (QCB) attacks.
By Aoying Zheng, Anqi Du, Zizhuang Deng, Yuxuan Chen
CS-Guard is a new benchmark that systematically evaluates guardrails for code generation security, covering 1,000 malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA) for text-to-code generation, as well as 331 code prompts for code-to-code generation. The study empirically tests nine guardrails across seven large language models, finding that many guardrails fail to prevent malicious code generation, with attack success rates reaching about 50% for text-to-code and up to nearly 100% for code-to-code and FSA scenarios. CS-Guard introduces a modular three-layer guardrail taxonomy and releases its benchmark and data to support future research.
By Jinyang Li, Mingyu Guo, Hung X. Nguyen
AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.
By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv:2508. 10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks.
By Jinhwa Kim, Ian G. Harris
arXiv:2607. 01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching.
By Joshua Adrian Cahyono
arXiv:2606.04483v2 Announce Type: replace
Abstract: Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the broader fa...
By Zhongze Luo, Ruihe Shi, Zhenshuai Yin, Haoyue Liu, Weixuan Wan, Xiaoying Tang
The paper introduces CodeScan, a black-box, vulnerability-oriented scanning framework designed to detect data poisoning and backdoor attacks in code generation large language models (LLMs). CodeScan operates by analyzing structural similarities across multiple code generations, normalizing them with abstract syntax tree (AST) techniques, and then applying LLM-based vulnerability analysis to identify recurring insecure patterns. Evaluations on 117 models across three architectures and multiple sizes show over 97% detection accuracy with fewer false positives compared to prior methods.
By Shenao Yan, Shan Jin, Shimaa Ahmed, Sunpreet Singh Arora, Yiwei Cai, Yizhen Wang, Yuan Hong