arXiv:2607. 26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate.
By Haoyu Zhang, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Mohammad Zandsalimy, Shanu Sushmita
A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. We show it can be breached by composing two attacks that are individually harmless against it: an established code-completion encoding and an established best-of-N search, neither of which exceeds 4.
The paper introduces Meta-Adaptive Multimodal Jailbreaking (MAMJ), a method that jointly optimizes an attack strategy prompt and attacker weights to generate more effective jailbreaks against vision‑language models. Using an LLM‑based critique to refine the strategy and group‑level success‑rate rewards to update the weights, MAMJ achieves high attack success rates on MM‑SafetyBench, outperforming existing baselines by up to 24.1 percentage points. The learned attacker also transfers to unseen models and remains robust against typical defenses, highlighting a systemic vulnerability in current VLMs.
By Benlei Cui, Shen Pang, Yuke Wang, Xuemei Dong, Yuwen Zhai, Jingqun Tang, Haiyang Yu, Hui Xue, Longtao Huang, Haiwen Hong
arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Yi Feng, Xiao Luo, Zijian Xiao, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
The paper presents a preprocessor that recovers and decodes encoded content in vision‑language models to close the decode gap that allows harmful requests to bypass safety classifiers. Evaluated against eleven encoding attacks, the preprocessor raises block rates from 0 % to 67‑90 % but also increases benign over‑refusal, and no configuration achieves an ensemble attack‑success rate below 40 % while keeping benign over‑refusal under 70 %. The study shows that closing one encoding channel merely relocates success rather than eliminating it, highlighting the limits of recovery‑based defenses.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng, Haowen Xu, Xiangchen Guan, Yang Chen, Zijian Xiao, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2607. 26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image...
arXiv:2608. 09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.
By Mingyu Luo, Ming Deng, Zilang Qiu, Yiming Cheng, Ci Tao, Xue Tan, Sijin Sun, Yangfu Li, Ping Chen, Jun Dai, Xiaoyan Sun
arXiv:2605. 03226v2 Announce Type: replace-cross Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset.
By Prakhar Gupta, Garv Shah, Donghua Zhang
arXiv:2606.04483v2 Announce Type: replace
Abstract: Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the broader fa...
By Zhongze Luo, Ruihe Shi, Zhenshuai Yin, Haoyue Liu, Weixuan Wan, Xiaoying Tang
arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.
By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka