arXiv Computer Vision

RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection

The paper introduces RED (Reconstruction Evolution Dynamics), a new framework for detecting AI-generated images that leverages the evolution of intermediate reconstruction stages rather than relying solely on static representations or endpoint discrepancies. RED uses a frozen multiscale VQ‑VAE and a frozen CLIP encoder to capture a reconstruction trajectory, then learns image‑adaptive stage weights from token negative log‑likelihoods provided by a frozen VAR model. Experiments on six benchmarks show RED achieves the highest average accuracy (92.5%) and precision (97.5%) among evaluated methods, and it remains robust to common image degradations.

arXiv Computer Vision
Aug 25

Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection

The paper proposes Artifact-Complementary Expert Fusion (ACEF), a two‑stage framework that enhances AI‑generated image detection by combining two types of reconstruction artifacts—VAE/DDIM and SRGAN—into aligned synthetic negatives. ACEF first builds artifact‑specific experts using LoRA adaptation on a frozen backbone, then fuses their multi‑layer evidence with Layer‑wise Artifact‑Complementary Fusion (LACF) to mitigate conflicts between artifact manifolds. Experiments on 13 benchmarks show that this approach improves generalizability over existing state‑of‑the‑art methods.

By Yiheng Li, Yang Yang, Wenhao Wang, Zichang Tan, Zecheng Lin, Li Gao, Zhen Lei
arXiv Computer Vision
Sep 2

Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video Detection

The paper introduces RIFT, a forensic framework for detecting AI-generated videos by exploiting a cross‑scale coupling mismatch between macro‑level temporal dynamics and micro‑level pixel residuals. RIFT comprises a macro stream that models expected temporal evolution, a micro stream that probes residual patterns, and a coupling divergence module that quantifies their conditional dependency. Experiments on VidProM and GenVidBench show near‑perfect F1‑scores and robust performance across different encoders.

By Siyu Li, Jin Yang, Weiheng Liang
arXiv Computer Vision
Aug 31

FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization

FUSED is a new framework that jointly detects and localizes AI-generated inpainting by combining low-level forensic cues with high-level semantic features through a sparsely-gated Mixture-of-Experts architecture. It predicts both an image-level manipulation score and a pixel-level mask of the inpainted region. On the OpenSDID cross-generator benchmark, FUSED outperforms existing methods, especially on unseen generators, and transfers effectively to the AutoSplice and CocoGlide benchmarks, doubling localization performance.

By Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska
arXiv Computer Vision
4d ago

Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics

Forensic Twins introduces a Self‑Supervised Residual Learning (SSRL) framework that trains on real images only, using a frozen forensic residual extractor to generate two disjoint crops per image. The pretext task suppresses semantic content, focusing the model on the stationary fingerprint of the image acquisition pipeline, and achieves 56.61% accuracy in attributing AI‑generator sources, outperforming prior zero‑shot methods. When combined with an offline Gaussian Mixture Model, the approach reaches 97.99% AUC across 27 unseen AI generators, including GANs, diffusion models, and commercial systems.

By Javier Mu\~noz-Haro, Ruben Tolosana, Ruben Vera-Rodriguez, Aythami Morales, Julian Fierrez
arXiv AI
Jun 2

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

arXiv:2606. 00101v1 Announce Type: cross Abstract: With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security.

By Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng
arXiv Computer Vision
Sep 16

Unifying Semantic Priors and High-Frequency Traces: Enhancing V-JEPA with Mixture-of-Experts for Robust Synthetic Image Forensics

The paper introduces MoE-JEPA, a dual‑stream deepfake detection model that combines a V‑JEPA backbone with a Residual Mixture‑of‑Experts mechanism and a noise stream branch. It further incorporates a Gated Attention Multiple Instance Learning module to refine spatial semantic understanding. On the SID‑Set benchmark, MoE‑JEPA achieves a new state‑of‑the‑art accuracy of 95.54%, outperforming much larger models.

By Simone Teglia, Irene Amerini