arXiv Machine Learning

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

arXiv:2607. 21776v1 Announce Type: new Abstract: Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries that current image-based detectors consistently fail to identify.

arXiv Machine Learning
Sep 1

Data Diversity, Not Frequency Invariance: A Controlled and Self-Audited Study of Compression-Robust Deepfake Detection

The study challenges the prevailing belief that frequency-based features and compression-invariant learning are essential for robust deepfake detection. Using a controlled, pre‑registered protocol, a simple EfficientNet‑B0 trained on diverse multi‑quality data outperformed the more complex CAFRL model across all compression levels, with a 3.66 AUC point advantage at CRF 40. After identifying and correcting four experimental defects, the authors found that frequency features added no marginal benefit, while data diversity—particularly real constant‑rate‑factor variants—proved to be the key factor for robustness against H.264 re‑encoding.

By Abbas Aliyev, Samir Rustamov
arXiv AI
Sep 25

Band-Attention Modulation Network for Robust Face Forgery Detection

The paper introduces Band-Attention Modulation Network (BAM‑Net), a face forgery detection framework that learns fine‑grained, adaptive modulation of frequency bands in the Discrete Cosine Transform spectrogram. BAM‑Net dynamically reweights anti‑diagonal frequency bands to enhance forgery‑related spectral cues while suppressing irrelevant information, then fuses this modulated frequency data with spatial features using a lightweight backbone with distance‑decayed attention. Experiments on FaceForensics++, Celeb‑DF, and DFDC show that BAM‑Net achieves state‑of‑the‑art performance and strong generalization across datasets, compression levels, and manipulation types.

By Zhida Zhang, Wenkui Yang, Xinlei Ma, Qihang Fan, Jie Cao
arXiv Computer Vision
Sep 25

$\unicode{x1F493}$Heartian: Physiology-Aware Relightable Gaussian Head Avatar

The paper introduces Heartian, a physiology‑aware framework that augments Gaussian head avatars with cardiac‑cycle‑dependent albedo modulation, enabling the encoding of remote photoplethysmography (rPPG) signals. By supervising with synchronized contact PPG, the method models the cardiac waveform as a sum of two Gaussians and learns per‑frame spatial residuals via a lightweight MLP. Experiments on 152 stationary recordings from UBFC‑rPPG, PURE, and MMPD show heart‑rate estimation errors as low as 0.29 bpm MAE and 0.38 % MAPE, while preserving reconstruction quality with negligible PSNR loss.

By Xiaoyue Fan, Jose Echevarria, Akshay Paruchuri, Kaan Ak\c{s}it