arXiv Computer Vision By Zihao Liao, Sheng Hong, Yu Chen

Spatiality-Frequency Domain Video Forgery Detection System Based on ResNet-LSTM-CBAM and DCT Hybrid Network

Read the original on arXiv Computer Vision →

The paper introduces a video forgery detection system that fuses spatial and frequency-domain features using a ResNet‑LSTM backbone with a Convolutional Block Attention Module (CBAM) and a Discrete Cosine Transform (DCT) module. Experiments on multiple benchmark datasets show that this hybrid architecture outperforms existing methods in distinguishing authentic from manipulated videos. Ablation and comparative studies confirm the individual contributions of each component, highlighting the model’s effectiveness across diverse forgery scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 25

Band-Attention Modulation Network for Robust Face Forgery Detection

The paper introduces Band-Attention Modulation Network (BAM‑Net), a face forgery detection framework that learns fine‑grained, adaptive modulation of frequency bands in the Discrete Cosine Transform spectrogram. BAM‑Net dynamically reweights anti‑diagonal frequency bands to enhance forgery‑related spectral cues while suppressing irrelevant information, then fuses this modulated frequency data with spatial features using a lightweight backbone with distance‑decayed attention. Experiments on FaceForensics++, Celeb‑DF, and DFDC show that BAM‑Net achieves state‑of‑the‑art performance and strong generalization across datasets, compression levels, and manipulation types.

By Zhida Zhang, Wenkui Yang, Xinlei Ma, Qihang Fan, Jie Cao
arXiv Computer Vision
6d ago

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.

By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li