FUSED is a new framework that jointly detects and localizes AI-generated inpainting by combining low-level forensic cues with high-level semantic features through a sparsely-gated Mixture-of-Experts architecture. It predicts both an image-level manipulation score and a pixel-level mask of the inpainted region. On the OpenSDID cross-generator benchmark, FUSED outperforms existing methods, especially on unseen generators, and transfers effectively to the AutoSplice and CocoGlide benchmarks, doubling localization performance.
By Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska
arXiv:2605.16879v2 Announce Type: replace
Abstract: With the rapid evolution of synthetic media, Image Manipulation Localization (IML) has emerged as a critical component in multimedia forensics for...
By Yunfei Wang, Bo Du, Zhe Yang, Xin Liu, Zhiyu Lin, Tianxin Xu, Ji-Zhe Zhou
arXiv:2609.38251v1 Announce Type: cross
Abstract: The rapid evolution of image manipulation techniques has raised growing public security concerns. Existing Image Forgery Localization (IFL) methods c...
By Chenqi Kong, Song Xia, Anwei Luo, Peisong He, Alex C. Kot, Yuming Fang
arXiv:2502. 19716v3 Announce Type: replace-cross Abstract: Recent advances in visual generative models have enabled the creation of highly realistic, fully AI-generated images without relying on real source content.
By Qijie Xu, Can Wang, Jiawei Chen, Siwei Lyu, Defang Chen
arXiv:2608. 04840v1 Announce Type: cross Abstract: Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence.
By Jacob Arndt, Debvrat Varshney, Philipe Dias, Nivedita Nukavarapu
arXiv:2608.20929v1 Announce Type: new
Abstract: AI-generated image manipulation localization identifies edited pixels, but its OOD performance lags behind image-level detection partly because pixel s...
By Haozhen Yan, Siyuan Shan, Zijian Yu, Youqi Wang, Yan Hong, Jun Lan, Jianfu Zhang
The paper introduces MoE-JEPA, a dual‑stream deepfake detection model that combines a V‑JEPA backbone with a Residual Mixture‑of‑Experts mechanism and a noise stream branch. It further incorporates a Gated Attention Multiple Instance Learning module to refine spatial semantic understanding. On the SID‑Set benchmark, MoE‑JEPA achieves a new state‑of‑the‑art accuracy of 95.54%, outperforming much larger models.
By Simone Teglia, Irene Amerini
The paper introduces DeformView, a wide‑baseline multi‑view dataset with pixel‑level annotations of geometric inconsistencies, and evaluates existing multi‑view consistency‑scoring methods, finding they transfer poorly to forensic localization tasks. To address this, the authors propose DEFECt3R, a lightweight learning‑based classifier that leverages cross‑view feature relationships and hard negative supervision to localize inconsistencies at the pixel level, achieving better performance and fewer false positives. Ablation studies confirm the importance of feature representations and correspondence quality for localization.
By Xander Staelens, Alb\'eric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael
Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. Highly realistic synthetic imagery produced for malicious purposes (deepfakes) can have major consequences in the remote sensing domain, where this data is a fundamental source of information for science applications, planning, logistics, and monitoring.
The paper proposes a unified forensics framework that extends traditional binary image manipulation detection to a multiclass setting—distinguishing real, fully synthetic, and tampered images. It adds a segmentation branch for pixel‑level localization of tampered regions, achieving higher classification accuracy and IoU scores compared to recent benchmarks. The authors provide the implementation on GitHub for reproducibility.
By Annalisa Gallina, Marco Fiorucci, Marco Brigo, Federica Battisti, Lamberto Ballan
Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization proposes a new framework that improves deepfake detection and interpretability. The approach introduces Feature-robust Augmentation—diversified degradation-aware strategies combined with supervised contrastive learning and a mean-teacher architecture—to maintain accuracy on low-quality images. For explanations, it employs evidence-grounded preference optimization, guiding the model to focus on genuine manipulation traces by learning from chosen-rejected explanation pairs that omit evidence or inject irrelevant details. The method achieved first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge and is publicly available on GitHub.
By Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu
arXiv:2607. 18230v1 Announce Type: cross Abstract: Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts.
By Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed Elhagry, Salwa K. Al Khatib, Tianjun Yao, Yonina C. Eldar, Jing-Hao Xue, Hao Li, Salman Khan, Zhiqiang Shen