As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors as black boxes, while the latent forgery-discriminative knowledge inside them remains largely unexplored.
ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.
By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li
arXiv:2606. 00101v1 Announce Type: cross Abstract: With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security.
By Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng
arXiv:2608. 06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety.
By Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
arXiv:2606. 01843v1 Announce Type: cross Abstract: Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that fail to transfer to unseen manipulations.
By Yihui Wang, Yonghui Yang, Jilong Liu, Fengbin Zhu, Le Wu, Tat-Seng Chua
arXiv:2609.38251v1 Announce Type: cross
Abstract: The rapid evolution of image manipulation techniques has raised growing public security concerns. Existing Image Forgery Localization (IFL) methods c...
By Chenqi Kong, Song Xia, Anwei Luo, Peisong He, Alex C. Kot, Yuming Fang
FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.
By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury
arXiv:2609.39585v1 Announce Type: new
Abstract: AI video generators have not only become harder to detect but are used to generate a diverse set of scenarios from landscapes to street views to animal...
By Sidharth Shanu, Gautam Kumar, Tej Singh
The paper introduces RIFT, a forensic framework for detecting AI-generated videos by exploiting a cross‑scale coupling mismatch between macro‑level temporal dynamics and micro‑level pixel residuals. RIFT comprises a macro stream that models expected temporal evolution, a micro stream that probes residual patterns, and a coupling divergence module that quantifies their conditional dependency. Experiments on VidProM and GenVidBench show near‑perfect F1‑scores and robust performance across different encoders.
By Siyu Li, Jin Yang, Weiheng Liang
arXiv:2606. 15880v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have been increasingly adopted in forensics for their robust semantic understanding.
By Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen, Ke-Yue Zhang, Yue Zhou, Caiyong Piao, Bin Li, Taiping Yao, Bo Wang, Youchang Xiao, Shouhong Ding
arXiv:2605.31192v2 Announce Type: replace
Abstract: Generalizable deepfake detection requires complementary forensic and semantic visual evidence. Specialist encoders capture subtle manipulation trac...
By Benedikt Hopf, Zongwei Wu, Radu Timofte
FUSED is a new framework that jointly detects and localizes AI-generated inpainting by combining low-level forensic cues with high-level semantic features through a sparsely-gated Mixture-of-Experts architecture. It predicts both an image-level manipulation score and a pixel-level mask of the inpainted region. On the OpenSDID cross-generator benchmark, FUSED outperforms existing methods, especially on unseen generators, and transfers effectively to the AutoSplice and CocoGlide benchmarks, doubling localization performance.
By Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska