arXiv Computation and Language

Your Voice Cloning System is Secretly a Voice Anonymizer

The paper demonstrates that the multilingual voice cloning model XTTSv2 can be repurposed for speaker anonymization without retraining. By conditioning on a pseudo-speaker and using an iterative refinement strategy, the authors balance privacy and intelligibility, achieving near‑optimal privacy (EER ≈ 0.49) and competitive speech quality across seven European languages. The method outperforms dedicated anonymization baselines and requires no language‑specific training.

arXiv AI
Jul 14

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

arXiv:2607. 09767v1 Announce Type: cross Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech.

By Adrien Schneider (M-PSI), Kacper Zabkowski (M-PSI), Anderson Augusma (M-PSI), Fr\'ed\'erique Letu\'e (SAM, SVH), Maria Camila Pinzon (M-PSI), Dominique Vaufreydaz (M-PSI)
arXiv Machine Learning
5d ago

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

DiffAnon is a diffusion‑based voice anonymization method that uses classifier‑free guidance to give users continuous, inference‑time control over how much prosody is preserved. By refining acoustic detail over semantic embeddings from an RVQ codec, the model allows smooth interpolation between strong anonymization and high prosodic fidelity within a single architecture. This is the first framework to provide structured, interpolatable prosody control while maintaining competitive privacy and utility across different operating points.

By Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews, Philipp Koehn, Berrak Sisman
arXiv AI
Jun 24

ZONOS2 Technical Report

arXiv:2606. 24320v1 Announce Type: cross Abstract: We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity.

By Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close, Beren Millidge
arXiv Machine Learning
Jul 27

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

arXiv:2607. 22304v1 Announce Type: new Abstract: Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented.

By Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez, Cole Looney, Xiaoliang Wu, Alexandra Livia Georgescu, Stefano Goria
Hugging Face Trending Papers
Jul 24

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice cloning is one such augmentation approach, but is typically evaluated on speech intelligibility (WER) or speaker similarity (SS) rather than on downstream performance, and it remains unclear whether these preserve the paralinguistic signal such tasks depend on.