The paper introduces the ASCII Attack, a single‑turn, black‑box method that embeds a harmful request within ASCII art and presents it as artwork to a large language model. By framing the request as artistic critique, the model can provide operational details that a plain request would normally be refused. Experiments across eleven models and eight harm topics show that the attack succeeds in 62% of cases versus 42% for direct controls, with the most vulnerable model achieving a 93% success rate.
By Da Cheng Gu, Yifei Dong, Xinghao Yang, Yongshun Gong, Wei Liu
arXiv:2607. 00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies.
By Shei Pern Chua, Fangzhao Wu
arXiv:2607. 26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate.
By Haoyu Zhang, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Mohammad Zandsalimy, Shanu Sushmita
A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. We show it can be breached by composing two attacks that are individually harmless against it: an established code-completion encoding and an established best-of-N search, neither of which exceeds 4.
arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
arXiv:2608. 10279v1 Announce Type: cross Abstract: Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, whereas repeated semantic classification of partial text can be costly and unstable.
By Christopher M. Frost
The paper investigates how allocating a query budget to structural depth rather than surface variation improves jailbreak success against the SAGE self‑check defense. By using a best‑of‑N approach over a code‑completion encoding, the authors achieve 67%, 22%, and 15% success rates on three open‑weight targets—far exceeding the 4.7% and 3.0% rates of single‑draw encoding and character‑search methods. The study demonstrates that depth of encoding and breadth of variation independently undermine transform and gate defenses, and that repeated sampling can inflate perceived robustness.
By Haoyu Zhang, Hanwen Liu, Yang Chen, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Xiao Luo, Yi Feng, Haowen Xu, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2608. 01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR).
By Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2606. 08403v1 Announce Type: cross Abstract: Text-centered prompt-injection defenses assume that the malicious signal is visible in one of the inspected text views.
By Mudit Sinha, Sanika Chavan
The paper introduces ‘Fool’s Gold’, a defensive deception technique for open‑weight language models that hardens them against safety‑removal attacks. By training decoy responses within a differentiable simulation of the attack, the method poisons the payoff of stripped refusal mechanisms, producing confident but falsified answers to hazardous requests while preserving benign behavior. Experiments on seven models (9B‑122B) show that 51‑90% of attacked‑state responses become decoys, with the defense accounting for 27‑84% of this effect, and that the defended 122B model remains within benign‑behavior budgets.
By Mark Russinovich
arXiv:2510.17904v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) are widely used because they process structures, syntax and code well, but this same ability also makes them par...
By Amirkia Rafiei Oskooei, Mehmet S. Aktas
Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.
By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov