The paper presents a preprocessor that recovers and decodes encoded content in vision‑language models to close the decode gap that allows harmful requests to bypass safety classifiers. Evaluated against eleven encoding attacks, the preprocessor raises block rates from 0 % to 67‑90 % but also increases benign over‑refusal, and no configuration achieves an ensemble attack‑success rate below 40 % while keeping benign over‑refusal under 70 %. The study shows that closing one encoding channel merely relocates success rather than eliminating it, highlighting the limits of recovery‑based defenses.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng, Haowen Xu, Xiangchen Guan, Yang Chen, Zijian Xiao, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Yi Feng, Xiao Luo, Zijian Xiao, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2607. 26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2605. 25889v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models reach high success rates on clean inputs but collapse under small adversarial perturbations: a $16/255$ PGD attack drops OpenVLA-7B's LIBERO success from $95\%$ to under $5\%$.
By Jianwei Tai
The paper introduces Meta-Adaptive Multimodal Jailbreaking (MAMJ), a method that jointly optimizes an attack strategy prompt and attacker weights to generate more effective jailbreaks against vision‑language models. Using an LLM‑based critique to refine the strategy and group‑level success‑rate rewards to update the weights, MAMJ achieves high attack success rates on MM‑SafetyBench, outperforming existing baselines by up to 24.1 percentage points. The learned attacker also transfers to unseen models and remains robust against typical defenses, highlighting a systemic vulnerability in current VLMs.
By Benlei Cui, Shen Pang, Yuke Wang, Xuemei Dong, Yuwen Zhai, Jingqun Tang, Haiyang Yu, Hui Xue, Longtao Huang, Haiwen Hong
arXiv:2606. 29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists.
By Subhadip Mitra
Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no defense, static steering, CAST, AlphaSteer, probe-gated) across seven instruction-tuned models (7-31B) and five attack types (GCG, AutoDAN, DeepInception, prefilling, intent laundering).
arXiv:2607.27910v2 Announce Type: replace
Abstract: Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder...
By Xiangyu Yin, Tora Bodin, Rohan Menon, Chih-Hong Cheng
arXiv:2607. 15207v1 Announce Type: new Abstract: World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction.
By Qi Li, Xingyi Yang, Xinchao Wang
The paper introduces FailBank, a four‑stage self‑evolving framework that transforms runtime feedback from safety shields into lasting policy improvements for vision‑language‑action (VLA) models. By using a counterfactual correction teacher, outcome‑aware admission, and guarded LoRA updates, FailBank converts useful shield proposals into corrective targets while preserving successful actions as anchors. Experiments on the VLA‑Arena benchmark show that FailBank boosts task success rates by up to 8.5 percentage points and reduces cumulative policy cost by up to 35.6%, outperforming both base policies and traditional runtime shielding.
By Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang
Guardrail models, which screen malicious prompts in LLM services, often use lightweight Transformers with short context windows and bucketed positional encodings. The study identifies a new failure mode called Overflip, where repeating a prompt causes the guardrail’s prediction to flip from malicious to benign as the sequence length increases. Experiments on nine popular guardrails show that 5 models exhibit MAL→BEN flips on 100 prompts, with flip rates ranging from 8% to 92% and first flips occurring between 2.6k and 9.4k tokens, highlighting a gradual attention dispersion distinct from traditional attention‑dilution attacks.
By Xu He, Chih-Hsuan Lin, Hung-Mao Chen, Junjie Xiong, Yan Zhai, Kun Sun
The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image...