arXiv Computer Vision
Aug 25

Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection

The paper proposes Artifact-Complementary Expert Fusion (ACEF), a two‑stage framework that enhances AI‑generated image detection by combining two types of reconstruction artifacts—VAE/DDIM and SRGAN—into aligned synthetic negatives. ACEF first builds artifact‑specific experts using LoRA adaptation on a frozen backbone, then fuses their multi‑layer evidence with Layer‑wise Artifact‑Complementary Fusion (LACF) to mitigate conflicts between artifact manifolds. Experiments on 13 benchmarks show that this approach improves generalizability over existing state‑of‑the‑art methods.

By Yiheng Li, Yang Yang, Wenhao Wang, Zichang Tan, Zecheng Lin, Li Gao, Zhen Lei
arXiv Computer Vision
6d ago

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.

By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li
arXiv Computer Vision
4d ago

RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection

The paper introduces RED (Reconstruction Evolution Dynamics), a new framework for detecting AI-generated images that leverages the evolution of intermediate reconstruction stages rather than relying solely on static representations or endpoint discrepancies. RED uses a frozen multiscale VQ‑VAE and a frozen CLIP encoder to capture a reconstruction trajectory, then learns image‑adaptive stage weights from token negative log‑likelihoods provided by a frozen VAR model. Experiments on six benchmarks show RED achieves the highest average accuracy (92.5%) and precision (97.5%) among evaluated methods, and it remains robust to common image degradations.

By Wenpeng Mu, Junshan Jin, Tanfeng Sun, Xinghao Jiang, Qiang Xu