arXiv AI

Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking

arXiv AI
Aug 12

MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

arXiv:2608. 10166v1 Announce Type: cross Abstract: Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored.

By Jie Cao, Qi Li, Zelin Zhang, Xiaodong Wu, Lingshuang Liu, Xiangman Li, Jianbing Ni
arXiv AI
2d ago

WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks

arXiv:2609.40031v1 Announce Type: cross Abstract: Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated co...

By Khaled Abud, Aleksey Yakushev, Aleksandr Akimenkov, Irina Serzhenko, Kirill Aistov, Egor Kovalev, Dmitry Obydenkov, Sergey Lavrushkin, Anastasia Antsiferova, Dmitriy Vatolin, Yury Markin, Kirill Lukianov
Hugging Face Trending Papers
Aug 10

DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models

The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infringement and misinformation. Although existing frequency-domain watermarking methods embed handcrafted geometric patterns into the initial latent noise prior to generation, they suffer from limited capacity and rigid pattern designs.

arXiv Computer Vision
5d ago

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

FeatMark is a watermarking framework that protects images from text‑to‑image diffusion model mimicry attacks by embedding small, scene‑consistent micro‑features instead of pixel‑level perturbations. It constructs domain‑specific feature banks, selects executable features, and injects them via mask‑guided concept editing to create highly localized, natural edits. Experiments on VGGFace2, CelebA‑HQ, and WikiArt show FeatMark remains robust against ten strong watermark removal attacks and several adaptive attacks, with minimal impact on perceptual quality and extending to video mimicry scenarios.

By Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao, Jason Xue, Haibo Hu
arXiv AI
Sep 2

One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

The paper introduces the concept of watermark laundering, where an attacker uses a single reconstruction prompt on public foundation image models to produce a visually faithful output that renders invisible watermarks undecodable. The authors evaluate this failure mode across six OpenAI and Google image editing models, three watermarking schemes, and 1,800 reconstructions, finding that OpenAI models cause the strongest payload disruption while Nano Banana 2 shows vulnerability of DwtDct under high-fidelity reconstruction. Prompt ablation experiments reveal that the disruption is driven by the reconstruction pathway itself rather than any specific removal instruction, highlighting prompt-conditioned reconstruction as a distinct attack interface.

By Jidong Yang, Qi Li, Wei Zong, Yang-Wai Chow, Willy Susilo, Huaike Yu, Chunpeng Wang, Suo Gao
Hugging Face Trending Papers
Sep 8

DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory

DRIFT is a black‑box attack that removes diffusion watermarks by combining partial forward diffusion with stochastic reverse resampling. It limits the source information available to a fixed‑depth recovery pipeline and uses stochastic reversal to explore alternative noise‑driven paths, refining fidelity only on updates rejected by the same verifier. Across nine watermarks, DRIFT achieves 98–100% attack success and the best image quality without requiring secret keys, verifier internals, or per‑image gradient optimization.

arXiv Machine Learning
Sep 10

DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory

DRIFT is a black‑box attack that removes diffusion watermarks by deflecting the generative trajectory. It combines partial forward diffusion with stochastic reverse resampling to limit the source information available to a fixed‑depth recovery pipeline and to explore alternative noise‑driven paths. Across nine watermarks, DRIFT achieves 98–100% success while preserving image quality, without requiring secret keys, verifier internals, or per‑image gradient optimization.

By Rui Bao, Zheng Gao, Xiaoyu Li, Xiaoyan Feng, Yang Song, Jiaojiao Jiang
Hugging Face Trending Papers
Jul 30

SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack

Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to capture the long-range dependencies of globally distributed watermark signals and resulting in an unfavorable trade-off between removal effectiveness and visual fidelity. In this paper, we propose SPFM-Net, a semantic-prior-guided and frequency-constrained Mamba framework for invisible watermark attack.

arXiv AI
Jun 10

Bypassing Copyright Protection in Diffusion-based Customization via Two-Stage Latent Feature Optimization

arXiv:2606. 09909v1 Announce Type: cross Abstract: With the growing concerns over copyright infringement in diffusion-based customization, adversarial attacks have emerged as a prominent defense strategy to prevent malicious content forgery in personalized image generation.

By Ziang Xu, Wenbo Yu, Hongyao Yu, Hao Fang, Jiawei Kong, Bin Chen, Hao Wu, Shu-Tao Xia, Zhiyong Wu