arXiv Computer Vision

Specialist-Generalist Fusion with Outcome-Supervised Rationales for Deepfake Detection

arXiv AI
Sep 18

FORGE: Forensic Reasoning with Grounded Evidence

FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.

By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury
arXiv AI
Aug 24

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization proposes a new framework that improves deepfake detection and interpretability. The approach introduces Feature-robust Augmentation—diversified degradation-aware strategies combined with supervised contrastive learning and a mean-teacher architecture—to maintain accuracy on low-quality images. For explanations, it employs evidence-grounded preference optimization, guiding the model to focus on genuine manipulation traces by learning from chosen-rejected explanation pairs that omit evidence or inject irrelevant details. The method achieved first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge and is publicly available on GitHub.

By Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu
arXiv Computer Vision
3d ago

MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection

The MSU team presents a modular approach to the Explainable Deepfake Detection Challenge, combining multiple DINOv3 backbones with Mesorch manipulation-localization features for real/fake classification. They incorporate a Grounding-DINO-based pseudo-mask pipeline to generate artifact evidence maps and a local contrastive objective to separate artifact from authenticity cues. For explanations, class-conditional Qwen3-VL models produce complex descriptions, which are then simplified by a GRPO-optimized text model, achieving high detection and explanation scores on the XPlainVerse dataset.

By Artem Filippov, Aleksandr Gushchin, Kirill Koltsov, Dmitriy Vatolin, Anastasia Antsiferova
arXiv Computer Vision
6d ago

Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling

This paper presents an interpretable deepfake detection framework that explicitly encodes physically grounded forensic cues to analyze spatially and temporally coherent facial features in video sequences. The method transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame with 68 structured descriptors across photometric, textural, geometric, and compression domains. These descriptors are processed by an LSTM to capture temporal dependencies, achieving strong F1-scores on four benchmark datasets and demonstrating robust cross-dataset generalization.

By Chahira Benhama, Mohand Sa\"id Allili, Assia Hamadene
arXiv AI
Jun 26

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

arXiv:2606. 26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images.

By Yangjun Wu, Keyu Yan, Yu Liu, Jingren Zhou, Fei Huang, Rong Zhang, Zhou Zhao, Fei Wu
arXiv Computer Vision
Sep 28

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.

By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li