The paper introduces Band-Attention Modulation Network (BAM‑Net), a face forgery detection framework that learns fine‑grained, adaptive modulation of frequency bands in the Discrete Cosine Transform spectrogram. BAM‑Net dynamically reweights anti‑diagonal frequency bands to enhance forgery‑related spectral cues while suppressing irrelevant information, then fuses this modulated frequency data with spatial features using a lightweight backbone with distance‑decayed attention. Experiments on FaceForensics++, Celeb‑DF, and DFDC show that BAM‑Net achieves state‑of‑the‑art performance and strong generalization across datasets, compression levels, and manipulation types.
By Zhida Zhang, Wenkui Yang, Xinlei Ma, Qihang Fan, Jie Cao
The paper introduces a video forgery detection system that fuses spatial and frequency-domain features using a ResNet‑LSTM backbone with a Convolutional Block Attention Module (CBAM) and a Discrete Cosine Transform (DCT) module. Experiments on multiple benchmark datasets show that this hybrid architecture outperforms existing methods in distinguishing authentic from manipulated videos. Ablation and comparative studies confirm the individual contributions of each component, highlighting the model’s effectiveness across diverse forgery scenarios.
By Zihao Liao, Sheng Hong, Yu Chen
arXiv:2609.14437v1 Announce Type: cross
Abstract: Deepfake detection systems often exhibit significant performance degradation when deployed on unseen manipulation methods, limiting their reliability...
By Arya Pulkit, Aditya Ruhela, Akarshan Kapoor, Arnav Bhavsar
arXiv:2603.14005v2 Announce Type: replace
Abstract: To generalize deepfake detectors to future unseen forgeries, most existing methods attempt to simulate the dynamically evolving forgery types using...
By Ming-Hui Liu, Harry Cheng, Xin Luo, Xin-Shun Xu, Mohan S. Kankanhalli
arXiv:2609.12668v1 Announce Type: new
Abstract: Recent deepfake detection studies increasingly suggest remote photoplethysmography (rPPG) signals as an authenticity cue. However, existing benchmarks...
By Chenxi Yang, Yassine Ouzar, Larbi Boubchir
The paper introduces Modality‑Specific Frequency Distillation (MSFD), a continual learning framework for video deepfake detection that separates spatial, temporal, and spatiotemporal features in the frequency domain. By preserving each modality independently and applying a cross‑modality decorrelation loss, MSFD adapts to new forgery patterns while maintaining performance across diverse continual deepfake video scenarios. Experiments demonstrate that this approach outperforms state‑of‑the‑art methods in both adaptation and retention.
By Taehoon Kim, Jongwook Choi, Heejae Jo, Byungmin Park, Jongwon Choi
arXiv:2609.01511v1 Announce Type: new
Abstract: Face forgery detectors often achieve strong results on controlled benchmarks, but their reliability under realistic image degradations remains limited....
By Lucas Cunha, Lucas Sotomaior, Lucas Gasperin, Beatriz Caldas, Eduardo Pianovski, Rayson Laroca
The paper introduces MoE-JEPA, a dual‑stream deepfake detection model that combines a V‑JEPA backbone with a Residual Mixture‑of‑Experts mechanism and a noise stream branch. It further incorporates a Gated Attention Multiple Instance Learning module to refine spatial semantic understanding. On the SID‑Set benchmark, MoE‑JEPA achieves a new state‑of‑the‑art accuracy of 95.54%, outperforming much larger models.
By Simone Teglia, Irene Amerini
arXiv:2609.07670v1 Announce Type: cross
Abstract: The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, dee...
By Xuechao Zou, Yi Zhou, Kai Li, Shun Zhang, Yuhui Chen, Congyan Lang, Junliang Xing
arXiv:2609.26274v1 Announce Type: new
Abstract: The rapid evolution of generative AI (e.g., Sora, Hunyuan) makes it essential to develop effective detection strategies that can generalize across ever...
By S. Hong, X. Q. Wang, C. Zhang, J. C. Wang, P. X. Duan, Y. W. Wang
The paper introduces S$^3$F-Net, a dual‑branch network that fuses spatial and spectral representations for medical image classification. It combines a deep spatial CNN with a shallow spectral encoder, SpectraNet, which uses a learnable SpectralFilter layer to process the full Fourier spectrum efficiently. Evaluated on four medical imaging datasets, S$^3$F-Net consistently outperforms spatial‑only baselines, achieving state‑of‑the‑art accuracy on BRISC2025 and surpassing deeper models on the Chest X‑Ray Pneumonia dataset.
By Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan
arXiv:2511. 10806v1 Announce Type: cross Abstract: Image deblurring is vital in computer vision, aiming to recover sharp images from blurry ones caused by motion or camera shake.
By Syed Mumtahin Mahmud, Mahdi Mohd Hossain Noki, Prothito Shovon Majumder, Abdul Mohaimen Al Radi, Md. Haider Ali, Md. Mosaddek Khan