arXiv Machine Learning By Pamela Kirui, Cho Hyuk, Qingzhong Liu, Haodi Jiang

An Explainable FFT-Based Spatial-Frequency Fusion Framework for Deepfake Detection

Read the original on arXiv Machine Learning →

arXiv:2607. 17441v1 Announce Type: cross Abstract: Deepfake generation has raised growing concerns regarding digital media authenticity, misinformation, identity fraud, and public trust.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Band-Attention Modulation Network for Robust Face Forgery Detection

The paper introduces Band-Attention Modulation Network (BAM‑Net), a face forgery detection framework that learns fine‑grained, adaptive modulation of frequency bands in the Discrete Cosine Transform spectrogram. BAM‑Net dynamically reweights anti‑diagonal frequency bands to enhance forgery‑related spectral cues while suppressing irrelevant information, then fuses this modulated frequency data with spatial features using a lightweight backbone with distance‑decayed attention. Experiments on FaceForensics++, Celeb‑DF, and DFDC show that BAM‑Net achieves state‑of‑the‑art performance and strong generalization across datasets, compression levels, and manipulation types.

By Zhida Zhang, Wenkui Yang, Xinlei Ma, Qihang Fan, Jie Cao
arXiv Computer Vision
Sep 24

Spatiality-Frequency Domain Video Forgery Detection System Based on ResNet-LSTM-CBAM and DCT Hybrid Network

The paper introduces a video forgery detection system that fuses spatial and frequency-domain features using a ResNet‑LSTM backbone with a Convolutional Block Attention Module (CBAM) and a Discrete Cosine Transform (DCT) module. Experiments on multiple benchmark datasets show that this hybrid architecture outperforms existing methods in distinguishing authentic from manipulated videos. Ablation and comparative studies confirm the individual contributions of each component, highlighting the model’s effectiveness across diverse forgery scenarios.

By Zihao Liao, Sheng Hong, Yu Chen
arXiv Computer Vision
Sep 4

Preserving Knowledge across Space and Time for Continual Video Deepfake Detection

The paper introduces Modality‑Specific Frequency Distillation (MSFD), a continual learning framework for video deepfake detection that separates spatial, temporal, and spatiotemporal features in the frequency domain. By preserving each modality independently and applying a cross‑modality decorrelation loss, MSFD adapts to new forgery patterns while maintaining performance across diverse continual deepfake video scenarios. Experiments demonstrate that this approach outperforms state‑of‑the‑art methods in both adaptation and retention.

By Taehoon Kim, Jongwook Choi, Heejae Jo, Byungmin Park, Jongwon Choi