arXiv:2609.38899v1 Announce Type: cross
Abstract: Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit poli...
By Wenyu Chen, Li Wang, Chuanchao Zang, Xiangtao Meng, Xinyu Gao, Jianing Wang, Zheng Li, Shanqing Guo
arXiv:2609.31032v1 Announce Type: cross
Abstract: Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video gen...
By Tianmeng Fang, Jiancheng Wang, Chen Wang, Liming Wang, Wei Wang, Jiayang Liu, Xiaochun Cao
The paper introduces TempJail, a temporal jailbreak framework targeting image‑to‑video generation models. It exploits a newly identified vulnerability where unsafe semantics arise from the composition of frames over time, rather than from single‑frame violations. By decomposing malicious captions into visual conditions and temporal instructions, and by employing controlled latent perturbations and template rewriting, TempJail achieves a 23.3 % higher attack success rate than prior methods on several commercial models.
By Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang
arXiv:2608. 19737v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning.
By Ling Zhou, Yihao Huang, Jingling Sun, Zhiwen Tian, Yi Zeng, Qihe Liu, Shijie Zhou
Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVL...
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos.
arXiv:2602.08136v2 Announce Type: replace-cross
Abstract: Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic...
By Md Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta Mehnaz
arXiv:2506. 22666v3 Announce Type: replace-cross Abstract: The rise of API-only access to state-of-the-art LLMs highlights the need for effective black-box jailbreak methods to identify model vulnerabilities in real-world settings.
By Anamika Lochab, Lu Yan, Patrick Pynadath, Xiangyu Zhang, Ruqi Zhang
The paper introduces the Detection Surface, a geometric framework that maps the decision boundaries of heterogeneous safety filters in text‑to‑image models. Using this insight, the authors propose CRACK, a multi‑agent debate system that iteratively mutates prompts, diagnoses layer‑specific constraints, and refines attacks to bypass composite defenses. Experiments demonstrate that CRACK can achieve attack success rates up to 99.63% while using fewer queries and preserving semantic fidelity.
By Kaiyan Wen, Shijie Zhang, Lu Yu, Guangdong Bai
arXiv:2607. 17779v1 Announce Type: new Abstract: Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images.
By Dongdong Yang, Deyue Zhang, Zhao Liu, Zonghao Ying, Wenzhuo Xu, Jiankai Jin, Xiangzheng Zhang, Quanchen Zou
DiSCO is a zero‑shot, black‑box defense for text‑to‑image models that operates solely at the prompt level. It expands prompts with a distribution‑guided suffix using beam search and contrastive scoring against safe and unsafe image pools generated by the target model, iteratively refining until safe content is produced. The method improves safety on the I2P benchmark under various red‑teaming attacks, reducing attack success rates by 37.7% and 25.13% while preserving semantic fidelity and image coherence.
By Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
arXiv:2606. 02111v1 Announce Type: cross Abstract: As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse.
By Choongwon Kang, Seungjong Sun, Hyunmin Jun, Jang Hyun Kim