arXiv Computer Vision

Towards Generalized Image Manipulation Localization via Score-based Model

arXiv AI
Jul 21

Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

arXiv:2607. 18230v1 Announce Type: cross Abstract: Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts.

By Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed Elhagry, Salwa K. Al Khatib, Tianjun Yao, Yonina C. Eldar, Jing-Hao Xue, Hao Li, Salman Khan, Zhiqiang Shen
arXiv Computer Vision
4d ago

FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization

FUSED is a new framework that jointly detects and localizes AI-generated inpainting by combining low-level forensic cues with high-level semantic features through a sparsely-gated Mixture-of-Experts architecture. It predicts both an image-level manipulation score and a pixel-level mask of the inpainted region. On the OpenSDID cross-generator benchmark, FUSED outperforms existing methods, especially on unseen generators, and transfers effectively to the AutoSplice and CocoGlide benchmarks, doubling localization performance.

By Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska
arXiv Computer Vision
1d ago

TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors

The paper introduces TIGA, a source‑image‑free, training‑free attack that injects adversarial properties into a diffusion model’s sampling trajectory to evade black‑box AIGC forensic detectors. TIGA aggregates gradients from white‑box surrogate detectors to create a transferable prior, then uses anisotropic directional search with finite‑difference queries to estimate and stabilize directions for the DDIM trajectory, applying frequency‑domain reshaping to reduce artifacts. Experiments demonstrate strong black‑box attack performance, transferability, and robustness to post‑processing while maintaining high perceptual quality.

By Xia Du, Zhuosen Bao, Zheng Lin, Jizhe Zhou, Chi-man Pun, Jun Luo, Symeon Chatzinotas
arXiv AI
Jun 2

Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization

arXiv:2606. 02178v1 Announce Type: cross Abstract: Recent advancements in generative AI have led to image editing models capable of producing realistic forgeries that evade traditional image forgery localization methods, as these approaches depend on physical noise absent in synthetic data.

By Yiming Wang, Baiqi Wu, Qingming Li, Jiahao Chen, Tong Zhang, Shouling Ji
arXiv AI
3d ago

CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification

The paper introduces CLIPure, a method for building an adversarially robust zero‑shot image classifier by purifying inputs in the latent space of CLIP. It formulates purification risk using KL divergence between denoising and attack processes via bidirectional SDEs, and proposes two variants: CLIPure‑Diff, which uses a diffusion prior, and CLIPure‑Cos, which relies on cosine similarity. Experiments on CIFAR‑10, ImageNet, and 13 other datasets show significant robustness gains, raising state‑of‑the‑art performance from 71.7% to 91.1% on CIFAR‑10 and from 59.6% to 72.6% on ImageNet.

By Mingkun Zhang, Keping Bi, Wei Chen, Jiafeng Guo, Xueqi Cheng
arXiv AI
Jun 11

Diffusion-based Cumulative Adversarial Purification for Vision Language Models

arXiv:2506. 03933v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications.

By Jia Fu, Yongtao Wu, Yihang Chen, Kunyu Peng, Xiao Zhang, Volkan Cevher, Sepideh Pashami, Anders Holst
arXiv AI
Aug 12

MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

arXiv:2608. 10166v1 Announce Type: cross Abstract: Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored.

By Jie Cao, Qi Li, Zelin Zhang, Xiaodong Wu, Lingshuang Liu, Xiangman Li, Jianbing Ni