arXiv AI

The Plot Twist: Jailbreaking Unified Multimodal Models with a Three-Act NarrativeAttack

The paper introduces NarrativeAttack, a jailbreak framework that exploits unified multimodal models (UMMs) by embedding a malicious query within a self‑contained three‑act visual narrative. The attack uses the model’s own image generator to create setup and resolution images, hiding the malicious event as a hidden climax, and concludes with an image‑based guessing game that forces the model to select the relevant answer. Experiments demonstrate that NarrativeAttack outperforms previous methods, achieving up to 88.25% attack success rate on Gemini‑2.5‑Flash, revealing a significant safety vulnerability in UMMs.

arXiv AI
Aug 19

COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

The paper introduces COMIC, a reference‑aware safety gate designed for multimodal large language models (MLLMs). COMIC detects the operation requested by a user, identifies visual targets through OCR and open‑vocabulary proposals, and evaluates safety on explicit operation‑target pairs, using max‑risk aggregation and quality‑aware routing to decide whether to allow or block a request. Experiments on several open‑source MLLMs and jailbreak benchmarks show that COMIC improves robustness while maintaining benign utility and efficiency.

By Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie, Huayue Gu, Chenyu Wang, Honghui Xu
arXiv AI
3d ago

CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition

CollageAttack is a black‑box jailbreak that exploits cross‑modal alignment flaws in text‑to‑image models by shifting harmful semantics into the image plane. It combines context‑relevant scenes, scene‑grounded textual carriers, and spatially distributed text fragments to produce images that reveal hidden harmful meaning. Experiments on both open‑weight and commercial models show success rates up to 86.0%, outperforming the strongest baseline by 18.5 percentage points and consistently generating more harmful outputs while preserving the source intent.

By Zhiyi Mou, Yao Lu, Wangze Ni, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou, Kui Ren
arXiv AI
Aug 19

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

The paper introduces MemJack, a memory‑augmented multi‑agent framework that automatically generates jailbreak attacks on Vision‑Language Models (VLMs) using benign natural images as visual anchors. MemJack discovers visual anchors, camouflages them semantically, evaluates responses, repairs via reflection, and replans dynamically, forming a closed‑loop attack pipeline. The authors also create MemJack‑Bench, a dataset of over 113,000 interactive multimodal jailbreak trajectories, and show that MemJack achieves a 71.48% attack success rate against Qwen3‑VL‑Plus, reaching 90% under extended budgets, outperforming other baselines on natural‑image evaluation.

By Jianhao Chen, Haoyang Chen, Hanjie Zhao, Haozhe Liang, Zheng Wang, Tieyun Qian
arXiv AI
Jul 21

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

arXiv:2607. 17779v1 Announce Type: new Abstract: Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images.

By Dongdong Yang, Deyue Zhang, Zhao Liu, Zonghao Ying, Wenzhuo Xu, Jiankai Jin, Xiangzheng Zhang, Quanchen Zou
arXiv AI
Jun 29

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

arXiv:2602. 10179v2 Announce Type: replace-cross Abstract: Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-text prompts.

By Jiacheng Hou, Yining Sun, Ruochong Jin, Haochen Han, Fangming Liu, Wai Kin Victor Chan, Alex Jinpeng Wang
arXiv Computation and Language
Sep 17

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

The paper introduces JMLLM, a multimodal jailbreaking approach that targets text, visual, and auditory inputs to expose vulnerabilities in large language models. It also presents TriJail, a new dataset containing jailbreak prompts across all three modalities. Experiments on TriJail and AdvBench show higher attack success rates and lower time overhead compared to existing methods.

By Yanxu Mao, Peipei Liu, Tiehan Cui, Zhaoteng Yan, Congying Liu, Datao You