arXiv:2512. 16300v3 Announce Type: replace Abstract: Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language models (MLLMs) with high-level semantic knowledge.
By Fanrui Zhang, Qiang Zhang, Sizhuo Zhou, Jianwen Sun, Chuanhao Li, Jiaxin Ai, Yukang Feng, Yujie Zhang, Wenjie Li, Zizhen Li, Yifan Chang, Jiawei Liu, Kaipeng Zhang
arXiv:2608. 06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety.
By Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
arXiv:2606. 26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images.
By Yangjun Wu, Keyu Yan, Yu Liu, Jingren Zhou, Fei Huang, Rong Zhang, Zhou Zhao, Fei Wu
The paper investigates an agentic framework for open‑world fake image detection that combines specialist detectors with per‑detector triage, prompting, and conflict‑aware evidence arbitration. Experiments across six configurations and three multimodal large language model backbones reveal that naive detector fusion yields high false‑positive rates, while triage and prompting consistently filter unreliable evidence. The most significant improvement comes from the reasoning component: a stronger judge markedly outperforms a weaker one, especially under distribution shift, and overall manipulation recall is nearly saturated, highlighting that the key challenge lies in calibrating trust and arbitrating conflicting forensic evidence rather than detecting manipulations themselves.
By Xianlong Li (IMT School for Advanced Studies Lucca, Italy), Pietro Bongini (University of Siena, Italy), Niccol\'o Pancino (University of Siena, Italy), Marco Blanchini (IMT School for Advanced Studies Lucca, Italy), Benedetta Tondi (University of Siena, Italy), Mauro Barni (University of Siena, Italy)
FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.
By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury
arXiv:2604. 02694v2 Announce Type: replace-cross Abstract: The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safety.
By Fanwei Zeng, Changtao Miao, Jing Huang, Zhiya Tan, Shutao Gong, Xiaoming Yu, Yang Wang, Weibin Yao, Joey Tianyi Zhou, Jianshu Li, Ying Yan
arXiv:2608. 08009v1 Announce Type: cross Abstract: Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4).
By Yichun Yeh, Yiheng Li, Xiaobo Hu, Zhen Lei, Yang Yang
arXiv:2609.36145v1 Announce Type: new
Abstract: Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective...
By Kaiqing Lin, Songze Li, Shen Chen, Yunfei Guo, Xiaoye Qiu, Haodong Li, Taiping Yao, Bo Wang, Youchang Xiao, Bin Li, Shouhong Ding
ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.
By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li
The paper introduces an evidence‑guided detector‑localizer‑reasoner system for text‑centric image forensics, addressing the challenges posed by AI‑generated content. It combines an image‑level authenticity detector, a localizer that extracts tampered regions, and an MLLM‑based reasoner that generates structured forensic reports grounded in the detected evidence. The system employs iterative difficulty‑aware mining and report‑mask consistency post‑processing, achieving a score of 0.638 and ranking second in the ACM Multimedia 2026 GenText‑Forensics Challenge.
By Peifeng Liu, Bin Li, Qingsong Zhang, Yangxin Yu, Leqing Chen, Xiaoye Qiu
Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization proposes a new framework that improves deepfake detection and interpretability. The approach introduces Feature-robust Augmentation—diversified degradation-aware strategies combined with supervised contrastive learning and a mean-teacher architecture—to maintain accuracy on low-quality images. For explanations, it employs evidence-grounded preference optimization, guiding the model to focus on genuine manipulation traces by learning from chosen-rejected explanation pairs that omit evidence or inject irrelevant details. The method achieved first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge and is publicly available on GitHub.
By Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu
arXiv:2609.05876v1 Announce Type: cross
Abstract: AI-generated image detection is increasingly open-world: new image generators produce highly realistic images that make visual artifacts harder to id...
By Jiaoyang Chen, Bin Hu, Jingyu Hu, Kun Zhou, Qin Zhang, Zhengzhe Liu