arXiv AI

Band-Attention Modulation Network for Robust Face Forgery Detection

The paper introduces Band-Attention Modulation Network (BAM‑Net), a face forgery detection framework that learns fine‑grained, adaptive modulation of frequency bands in the Discrete Cosine Transform spectrogram. BAM‑Net dynamically reweights anti‑diagonal frequency bands to enhance forgery‑related spectral cues while suppressing irrelevant information, then fuses this modulated frequency data with spatial features using a lightweight backbone with distance‑decayed attention. Experiments on FaceForensics++, Celeb‑DF, and DFDC show that BAM‑Net achieves state‑of‑the‑art performance and strong generalization across datasets, compression levels, and manipulation types.

arXiv Computer Vision
Sep 24

Spatiality-Frequency Domain Video Forgery Detection System Based on ResNet-LSTM-CBAM and DCT Hybrid Network

The paper introduces a video forgery detection system that fuses spatial and frequency-domain features using a ResNet‑LSTM backbone with a Convolutional Block Attention Module (CBAM) and a Discrete Cosine Transform (DCT) module. Experiments on multiple benchmark datasets show that this hybrid architecture outperforms existing methods in distinguishing authentic from manipulated videos. Ablation and comparative studies confirm the individual contributions of each component, highlighting the model’s effectiveness across diverse forgery scenarios.

By Zihao Liao, Sheng Hong, Yu Chen
arXiv Computer Vision
Sep 17

Generalizable Face Forgery Detection via Separable Prompt Learning

The paper introduces Separable Prompt Learning (SePL), a method that enhances face forgery detection by leveraging CLIP’s textual encoder through two separate learnable prompts. SePL incorporates a cross-modality alignment strategy and specific objectives to distill forgery knowledge from CLIP. Experiments show that SePL outperforms existing approaches in cross-dataset and cross-method evaluations.

By Enrui Yang, Baoyuan Wu, Yuezun Li
arXiv Computer Vision
Sep 18

IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models

The paper introduces IMFD, an end‑to‑end multi‑face forgery detector that uses instruction‑based Large Vision‑Language Models (LVLMs). IMFD jointly localizes faces and predicts forgery labels in a single stage, explicitly incorporating predicted face bounding boxes into the textual instruction to improve grounding and detection. Experiments on converted multi‑face forgery datasets show that IMFD outperforms several state‑of‑the‑art methods.

By Dasom Choi, Sangjun Moon, Hyeongchan Im, Jaeeon Park, Jingun Kwon, Hidetaka Kamigaito, Taro Watanabe, Manabu Okumura