arXiv:2607. 17279v1 Announce Type: cross Abstract: Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks.
By Xingkai Peng, Jun Jiang, Jiayang Liu, Kejiang Chen, Weiming Zhang
arXiv:2609.38899v1 Announce Type: cross
Abstract: Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit poli...
By Wenyu Chen, Li Wang, Chuanchao Zang, Xiangtao Meng, Xinyu Gao, Jianing Wang, Zheng Li, Shanqing Guo
AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.
By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv:2606. 05609v1 Announce Type: cross Abstract: As large language models (LLMs) are widely deployed, identifying their vulnerability through jailbreak attacks becomes increasingly critical.
By Seungwon Jeong, Jiwoo Jeong, Hyeonjin Kim, Yunseok Lee, Woojin Lee
arXiv:2608. 19737v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning.
By Ling Zhou, Yihao Huang, Jingling Sun, Zhiwen Tian, Yi Zeng, Qihe Liu, Shijie Zhou
The paper introduces Fair-ASR, a protocol for evaluating black‑box jailbreak attacks using a shared target‑call budget, addressing the bias of prior studies that rely solely on attack success rate. Re‑evaluating 11 attacks under this protocol shows significant shifts in rankings and highlights that many methods are not efficient in both target and attacker calls. The authors then present ReCode, a budget‑efficient attack that combines desensitization rewriting with low‑cost primitives, achieving 85% ASR on GPT‑5 with only 7.19 attacker calls per request under a 20‑target‑call budget.
By Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang
The paper introduces TempJail, a temporal jailbreak framework targeting image‑to‑video generation models. It exploits a newly identified vulnerability where unsafe semantics arise from the composition of frames over time, rather than from single‑frame violations. By decomposing malicious captions into visual conditions and temporal instructions, and by employing controlled latent perturbations and template rewriting, TempJail achieves a 23.3 % higher attack success rate than prior methods on several commercial models.
By Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang
The paper introduces TACS, a trajectory‑aware candidate selection framework designed to improve jailbreak suffix optimization for large language models. Traditional gradient‑based methods choose candidates based solely on the lowest current loss, which the authors argue is myopic and can lead to reward hacking. TACS augments per‑step evaluation with a trajectory‑aware proxy, reference‑policy regularization, and a chi‑squared correction to encourage selections that remain effective beyond the immediate step. Experiments on HarmBench show that TACS consistently outperforms strong baselines, achieving higher attack success rates and more stable optimization behavior.
By Shiliang Xiao
Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVL...
The paper introduces Fair-ASR, a new evaluation protocol for black-box jailbreak attacks that uses shared target-call budgets to provide a fair comparison across methods. Re‑evaluating 11 attacks under this protocol shows that rankings shift significantly with different budgets, and that simple perturbations and templates remain competitive. The authors also present ReCode, a budget‑efficient attack that combines desensitization rewriting with low‑cost primitives, achieving 85% ASR on GPT‑5 with only 7.19 attacker calls per request under a 20‑call budget.
The paper investigates how allocating a query budget to structural depth rather than surface variation improves jailbreak success against the SAGE self‑check defense. By using a best‑of‑N approach over a code‑completion encoding, the authors achieve 67%, 22%, and 15% success rates on three open‑weight targets—far exceeding the 4.7% and 3.0% rates of single‑draw encoding and character‑search methods. The study demonstrates that depth of encoding and breadth of variation independently undermine transform and gate defenses, and that repeated sampling can inflate perceived robustness.
By Haoyu Zhang, Hanwen Liu, Yang Chen, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Xiao Luo, Yi Feng, Haowen Xu, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2608. 03070v1 Announce Type: cross Abstract: Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much protection these safeguards provide, or how consistently across developers.
By Jasper Timm, Lukas Struppek, Ziwei Xu, Grace Cheong, Oscar Mata, Dan Zhao, Mick Yang, Isadora De Andrade, Xiaojun Jia, Yiming Li, Samuel Bauer, Heather McIntyre, Adam Gleave, Edward Yee, Kellin Pelrine