arXiv AI

Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering

arXiv:2606. 17257v1 Announce Type: cross Abstract: Open-weight video diffusion models can generate photorealistic unsafe content, from violence to misinformation, yet existing defenses either require expensive safety fine-tuning that degrades general capability, or apply external filters that are trivially bypassed by adversarial prompts.

arXiv AI
Jun 6

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

arXiv:2606. 05290v1 Announce Type: cross Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture.

By Tobia Poppi, Silvia Cappelletti, Sara Sarto, Florian Schiffers, Garin Kessler, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Hugging Face Trending Papers
Jul 7

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization.

arXiv Computer Vision
Sep 4

SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models

SafeRI proposes an on-demand safety alignment approach for large vision-language models, contrasting with existing always-on methods that globally modify model behavior. The framework uses a lightweight recognizer to evaluate token-level safety during autoregressive generation, gating a LoRA module that only activates when unsafe content is detected. By training the LoRA on unsafe prefixes and safe continuations, SafeRI redirects unsafe generations back to safe responses without perturbing the model’s original reasoning path.

By Caoyuan Ma, Tian Gu, Wenpu Liu, Weichu Xie, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Zheng Wang, Yinqiang Zheng
arXiv Computer Vision
Aug 27

DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

DEFUSE is a backdoor detection framework for self‑supervised encoders that uses a conditional diffusion generative model to estimate representation‑conditioned image likelihoods. By fine‑tuning a pretrained diffusion model, DEFUSE performs semantic reconstruction in a reference encoder’s representation space, enabling it to detect backdoors without needing uninfected data or precomputed pseudo‑labels. Experiments show that DEFUSE outperforms existing detectors on both visual SSL and vision‑language encoders, reducing reliance on prior knowledge of the victim model or attack strategy.

By Tuo Chen, Jie Gui, Minjing Dong, Lanting Fang, Ju Jia, Benlei Cui, Jian Liu
arXiv Computer Vision
Aug 28

TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models

The paper introduces TempJail, a temporal jailbreak framework targeting image‑to‑video generation models. It exploits a newly identified vulnerability where unsafe semantics arise from the composition of frames over time, rather than from single‑frame violations. By decomposing malicious captions into visual conditions and temporal instructions, and by employing controlled latent perturbations and template rewriting, TempJail achieves a 23.3 % higher attack success rate than prior methods on several commercial models.

By Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang
arXiv Computer Vision
Sep 11

A Multi-View and Confusion-Guided Ensemble Framework for Robust Synthetic Image Attribution

The paper introduces a multi‑view, confusion‑guided ensemble framework for synthetic image attribution, combining FFT‑ConvNeXt, DINOv2, CLIP, and Xception to capture frequency, semantic, and forensic cues. Extensive data augmentation simulates realistic post‑processing, while a binary expert classifier and class‑adaptive confidence calibration address ambiguities between similar diffusion models. The approach achieved 99.53% on the public leaderboard and 99.20% on the private leaderboard for the ICANN 2026 DLMMDD Workshop challenge.

By Zuomin Qu