Audio-Driven Adversarial Defense for 3D Talking Face Generation with totally Visual Fidelity Preservation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 03985v1 Announce Type: cross Abstract: Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS).
The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.
arXiv:2606. 11615v1 Announce Type: cross Abstract: The widespread adoption of face recognition (FR) technologies raises serious privacy concerns, as facial data can be exploited without consent.
The paper introduces a two‑stage speech anonymization framework that preserves both linguistic content and acoustic identity. It replaces personally identifiable information using a generative editing model and applies a flow‑matching anonymization technique (F3‑VA) to create diverse, distinct anonymized speakers. The authors evaluate privacy with speaker verification metrics and utility by training ASR, TTS, and SER models from scratch, showing stronger privacy protection with minimal utility loss compared to existing baselines.
arXiv:2609.17422v1 Announce Type: new Abstract: Audio-driven digital human generation plays an important role in virtual communication, immersive interaction, and media production. With the developme...
arXiv:2607. 26742v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.