arXiv AI

Anonymization, Not Elimination: Utility-Preserved Speech Anonymization

arXiv Machine Learning
5d ago

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

DiffAnon is a diffusion‑based voice anonymization method that uses classifier‑free guidance to give users continuous, inference‑time control over how much prosody is preserved. By refining acoustic detail over semantic embeddings from an RVQ codec, the model allows smooth interpolation between strong anonymization and high prosodic fidelity within a single architecture. This is the first framework to provide structured, interpolatable prosody control while maintaining competitive privacy and utility across different operating points.

By Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews, Philipp Koehn, Berrak Sisman
arXiv AI
Jul 14

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

arXiv:2607. 09767v1 Announce Type: cross Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech.

By Adrien Schneider (M-PSI), Kacper Zabkowski (M-PSI), Anderson Augusma (M-PSI), Fr\'ed\'erique Letu\'e (SAM, SVH), Maria Camila Pinzon (M-PSI), Dominique Vaufreydaz (M-PSI)
arXiv Computation and Language
Aug 28

Your Voice Cloning System is Secretly a Voice Anonymizer

The paper demonstrates that the multilingual voice cloning model XTTSv2 can be repurposed for speaker anonymization without retraining. By conditioning on a pseudo-speaker and using an iterative refinement strategy, the authors balance privacy and intelligibility, achieving near‑optimal privacy (EER ≈ 0.49) and competitive speech quality across seven European languages. The method outperforms dedicated anonymization baselines and requires no language‑specific training.

By Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu
arXiv AI
1d ago

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability

The paper introduces Traceable TTS, a framework that enables Text‑to‑Speech systems to attribute synthesized speech to their source models without embedding explicit watermarks. By jointly training the TTS model and a discriminator, the method improves traceability generalization while maintaining or slightly enhancing audio quality. This represents the first attempt at watermark‑free TTS with strong traceability, and the authors plan to release the code to support further research.

By Yuxiang Zhao, Yunchong Xiao, Yushen Chen, Zhikang Niu, Shuai Wang, Kai Yu, Xie Chen