Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 21776v1 Announce Type: new Abstract: Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries that current image-based detectors consistently fail to identify.
arXiv:2609.12668v1 Announce Type: new Abstract: Recent deepfake detection studies increasingly suggest remote photoplethysmography (rPPG) signals as an authenticity cue. However, existing benchmarks...
arXiv:2609.38019v1 Announce Type: new Abstract: We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rat...
arXiv:2608.23363v1 Announce Type: cross Abstract: Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deep...
Remote Photoplethysmography (rPPG) enables contactless pulse estimation from facial videos, serving as a vital tool for health monitoring. However, current deep learning methods often struggle under complex disturbances, particularly varying illumination, facial expressions, and unconstrained head movements.
The paper introduces AV-STE, a modular streaming audio‑visual front‑end that enhances corrupted semantic speech tokens using noisy audio and lip video before they reach a frozen speech LLM. By preserving the downstream dialogue model’s pretrained conversational abilities, AV‑STE improves response coherence from 1.42 to 1.91 in same‑dataset speaker interference scenarios while maintaining turn‑taking behavior. These gains also transfer to out‑of‑domain Seamless Interaction.