arXiv AI By Zhiyi Mou, Yao Lu, Wangze Ni, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou, Kui Ren

CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition

Read the original on arXiv AI →

CollageAttack is a black‑box jailbreak that exploits cross‑modal alignment flaws in text‑to‑image models by shifting harmful semantics into the image plane. It combines context‑relevant scenes, scene‑grounded textual carriers, and spatially distributed text fragments to produce images that reveal hidden harmful meaning. Experiments on both open‑weight and commercial models show success rates up to 86.0%, outperforming the strongest baseline by 18.5 percentage points and consistently generating more harmful outputs while preserving the source intent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

The Plot Twist: Jailbreaking Unified Multimodal Models with a Three-Act NarrativeAttack

The paper introduces NarrativeAttack, a jailbreak framework that exploits unified multimodal models (UMMs) by embedding a malicious query within a self‑contained three‑act visual narrative. The attack uses the model’s own image generator to create setup and resolution images, hiding the malicious event as a hidden climax, and concludes with an image‑based guessing game that forces the model to select the relevant answer. Experiments demonstrate that NarrativeAttack outperforms previous methods, achieving up to 88.25% attack success rate on Gemini‑2.5‑Flash, revealing a significant safety vulnerability in UMMs.

By Shaoxiong Guo, Tianyi Du, Lijun Li, Yuyao Wu, Jie Li, Jing Shao
Hugging Face Trending Papers
Jul 7

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization.

arXiv Computation and Language
Aug 25

Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

The paper introduces TA-SPA, a black‑box jailbreak method for multimodal large language models that generates transferable perturbations in a text‑anchored semantic space. It combines Text‑Anchored Semantic Factorization (TASF) to separate cross‑modal semantic factors from modality‑specific residuals with Semantic‑Preserving Augmentation (SPA) to diversify harmful target anchors while maintaining semantic consistency. Experiments demonstrate strong attack effectiveness and transferability to commercial MLLMs, with competitive performance against representative defenses.

By Wenyun Li, Guiping Cao, Xiangyuan Lan, Zheng Zhang
arXiv AI
Aug 19

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

DiSCO is a zero‑shot, black‑box defense for text‑to‑image models that operates solely at the prompt level. It expands prompts with a distribution‑guided suffix using beam search and contrastive scoring against safe and unsafe image pools generated by the target model, iteratively refining until safe content is produced. The method improves safety on the I2P benchmark under various red‑teaming attacks, reducing attack success rates by 37.7% and 25.13% while preserving semantic fidelity and image coherence.

By Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem