arXiv AI By Adeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, Mubarak Shah

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

Read the original on arXiv AI →

arXiv:2607. 00402v1 Announce Type: cross Abstract: Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 7

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization.

arXiv AI
Jun 24

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.

By Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu
arXiv AI
Sep 24

InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation

InGuard introduces an inner guardrail for text-to-image generation that operates within the model’s own representations, avoiding external classifiers. It grades prompts using the text encoder’s embeddings, modifies risky embeddings with SAGE to produce safe images, and employs a latent detector to halt generation early. Evaluated on the RevGen Safety Benchmark, InGuard achieves a 97.9–98.8% safety rate across five open-weight models while reducing benign disturbances, model parameters, and denoising steps.

By Zeyu Wang, Xiaodan Li, Zhiwen Li, Yuefeng Chen, Hui Xue
arXiv AI
3d ago

CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition

CollageAttack is a black‑box jailbreak that exploits cross‑modal alignment flaws in text‑to‑image models by shifting harmful semantics into the image plane. It combines context‑relevant scenes, scene‑grounded textual carriers, and spatially distributed text fragments to produce images that reveal hidden harmful meaning. Experiments on both open‑weight and commercial models show success rates up to 86.0%, outperforming the strongest baseline by 18.5 percentage points and consistently generating more harmful outputs while preserving the source intent.

By Zhiyi Mou, Yao Lu, Wangze Ni, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou, Kui Ren