arXiv:2608. 19737v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning.
By Ling Zhou, Yihao Huang, Jingling Sun, Zhiwen Tian, Yi Zeng, Qihe Liu, Shijie Zhou
Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVL...
arXiv:2607. 17279v1 Announce Type: cross Abstract: Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks.
By Xingkai Peng, Jun Jiang, Jiayang Liu, Kejiang Chen, Weiming Zhang
arXiv:2606. 02111v1 Announce Type: cross Abstract: As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse.
By Choongwon Kang, Seungjong Sun, Hyunmin Jun, Jang Hyun Kim
The paper introduces TempJail, a temporal jailbreak framework targeting image‑to‑video generation models. It exploits a newly identified vulnerability where unsafe semantics arise from the composition of frames over time, rather than from single‑frame violations. By decomposing malicious captions into visual conditions and temporal instructions, and by employing controlled latent perturbations and template rewriting, TempJail achieves a 23.3 % higher attack success rate than prior methods on several commercial models.
By Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang
arXiv:2503.06989v5 Announce Type: replace-cross
Abstract: Recently, Multimodal Large Language Models (MLLMs) have demonstrated their superior ability in understanding multimodal content. However, the...
By Wenzhuo Xu, Zhipeng Wei, Xiongtao Sun, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos.
arXiv:2609.31032v1 Announce Type: cross
Abstract: Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video gen...
By Tianmeng Fang, Jiancheng Wang, Chen Wang, Liming Wang, Wei Wang, Jiayang Liu, Xiaochun Cao
The paper introduces MemJack, a memory‑augmented multi‑agent framework that automatically generates jailbreak attacks on Vision‑Language Models (VLMs) using benign natural images as visual anchors. MemJack discovers visual anchors, camouflages them semantically, evaluates responses, repairs via reflection, and replans dynamically, forming a closed‑loop attack pipeline. The authors also create MemJack‑Bench, a dataset of over 113,000 interactive multimodal jailbreak trajectories, and show that MemJack achieves a 71.48% attack success rate against Qwen3‑VL‑Plus, reaching 90% under extended budgets, outperforming other baselines on natural‑image evaluation.
By Jianhao Chen, Haoyang Chen, Hanjie Zhao, Haozhe Liang, Zheng Wang, Tieyun Qian
The paper introduces the Detection Surface, a geometric framework that maps the decision boundaries of heterogeneous safety filters in text‑to‑image models. Using this insight, the authors propose CRACK, a multi‑agent debate system that iteratively mutates prompts, diagnoses layer‑specific constraints, and refines attacks to bypass composite defenses. Experiments demonstrate that CRACK can achieve attack success rates up to 99.63% while using fewer queries and preserving semantic fidelity.
By Kaiyan Wen, Shijie Zhang, Lu Yu, Guangdong Bai
The paper introduces JMLLM, a multimodal jailbreaking approach that targets text, visual, and auditory inputs to expose vulnerabilities in large language models. It also presents TriJail, a new dataset containing jailbreak prompts across all three modalities. Experiments on TriJail and AdvBench show higher attack success rates and lower time overhead compared to existing methods.
By Yanxu Mao, Peipei Liu, Tiehan Cui, Zhaoteng Yan, Congying Liu, Datao You
The paper introduces NarrativeAttack, a jailbreak framework that exploits unified multimodal models (UMMs) by embedding a malicious query within a self‑contained three‑act visual narrative. The attack uses the model’s own image generator to create setup and resolution images, hiding the malicious event as a hidden climax, and concludes with an image‑based guessing game that forces the model to select the relevant answer. Experiments demonstrate that NarrativeAttack outperforms previous methods, achieving up to 88.25% attack success rate on Gemini‑2.5‑Flash, revealing a significant safety vulnerability in UMMs.
By Shaoxiong Guo, Tianyi Du, Lijun Li, Yuyao Wu, Jie Li, Jing Shao