Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 03985v1 Announce Type: cross Abstract: Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS).
arXiv:2601.19956v2 Announce Type: replace-cross Abstract: As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge eme...
DiffAnon is a diffusion‑based voice anonymization method that uses classifier‑free guidance to give users continuous, inference‑time control over how much prosody is preserved. By refining acoustic detail over semantic embeddings from an RVQ codec, the model allows smooth interpolation between strong anonymization and high prosodic fidelity within a single architecture. This is the first framework to provide structured, interpolatable prosody control while maintaining competitive privacy and utility across different operating points.
arXiv:2607. 09767v1 Announce Type: cross Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech.
The paper demonstrates that the multilingual voice cloning model XTTSv2 can be repurposed for speaker anonymization without retraining. By conditioning on a pseudo-speaker and using an iterative refinement strategy, the authors balance privacy and intelligibility, achieving near‑optimal privacy (EER ≈ 0.49) and competitive speech quality across seven European languages. The method outperforms dedicated anonymization baselines and requires no language‑specific training.
arXiv:2606. 29897v1 Announce Type: cross Abstract: Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability.