arXiv AI

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

The paper introduces MemJack, a memory‑augmented multi‑agent framework that automatically generates jailbreak attacks on Vision‑Language Models (VLMs) using benign natural images as visual anchors. MemJack discovers visual anchors, camouflages them semantically, evaluates responses, repairs via reflection, and replans dynamically, forming a closed‑loop attack pipeline. The authors also create MemJack‑Bench, a dataset of over 113,000 interactive multimodal jailbreak trajectories, and show that MemJack achieves a 71.48% attack success rate against Qwen3‑VL‑Plus, reaching 90% under extended budgets, outperforming other baselines on natural‑image evaluation.

arXiv AI
Aug 19

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

DiSCO is a zero‑shot, black‑box defense for text‑to‑image models that operates solely at the prompt level. It expands prompts with a distribution‑guided suffix using beam search and contrastive scoring against safe and unsafe image pools generated by the target model, iteratively refining until safe content is produced. The method improves safety on the I2P benchmark under various red‑teaming attacks, reducing attack success rates by 37.7% and 25.13% while preserving semantic fidelity and image coherence.

By Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
arXiv AI
Jun 29

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

arXiv:2602. 10179v2 Announce Type: replace-cross Abstract: Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-text prompts.

By Jiacheng Hou, Yining Sun, Ruochong Jin, Haochen Han, Fangming Liu, Wai Kin Victor Chan, Alex Jinpeng Wang
arXiv Computer Vision
Aug 31

Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models

The paper introduces Meta-Adaptive Multimodal Jailbreaking (MAMJ), a method that jointly optimizes an attack strategy prompt and attacker weights to generate more effective jailbreaks against vision‑language models. Using an LLM‑based critique to refine the strategy and group‑level success‑rate rewards to update the weights, MAMJ achieves high attack success rates on MM‑SafetyBench, outperforming existing baselines by up to 24.1 percentage points. The learned attacker also transfers to unseen models and remains robust against typical defenses, highlighting a systemic vulnerability in current VLMs.

By Benlei Cui, Shen Pang, Yuke Wang, Xuemei Dong, Yuwen Zhai, Jingqun Tang, Haiyang Yu, Hui Xue, Longtao Huang, Haiwen Hong
Hugging Face Trending Papers
Jun 24

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos.

arXiv AI
6d ago

The Plot Twist: Jailbreaking Unified Multimodal Models with a Three-Act NarrativeAttack

The paper introduces NarrativeAttack, a jailbreak framework that exploits unified multimodal models (UMMs) by embedding a malicious query within a self‑contained three‑act visual narrative. The attack uses the model’s own image generator to create setup and resolution images, hiding the malicious event as a hidden climax, and concludes with an image‑based guessing game that forces the model to select the relevant answer. Experiments demonstrate that NarrativeAttack outperforms previous methods, achieving up to 88.25% attack success rate on Gemini‑2.5‑Flash, revealing a significant safety vulnerability in UMMs.

By Shaoxiong Guo, Tianyi Du, Lijun Li, Yuyao Wu, Jie Li, Jing Shao