arXiv AI

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization

arXiv:2607. 03985v1 Announce Type: cross Abstract: Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS).

arXiv AI
Sep 4

Anonymization, Not Elimination: Utility-Preserved Speech Anonymization

The paper introduces a two‑stage speech anonymization framework that preserves both linguistic content and acoustic identity. It replaces personally identifiable information using a generative editing model and applies a flow‑matching anonymization technique (F3‑VA) to create diverse, distinct anonymized speakers. The authors evaluate privacy with speaker verification metrics and utility by training ASR, TTS, and SER models from scratch, showing stronger privacy protection with minimal utility loss compared to existing baselines.

By Yunchong Xiao, Yuxiang Zhao, Ziyang Ma, Shuai Wang, Kai Yu, Jiachun Liao, Xie Chen
arXiv Computation and Language
Aug 28

Your Voice Cloning System is Secretly a Voice Anonymizer

The paper demonstrates that the multilingual voice cloning model XTTSv2 can be repurposed for speaker anonymization without retraining. By conditioning on a pseudo-speaker and using an iterative refinement strategy, the authors balance privacy and intelligibility, achieving near‑optimal privacy (EER ≈ 0.49) and competitive speech quality across seven European languages. The method outperforms dedicated anonymization baselines and requires no language‑specific training.

By Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu
arXiv AI
Jul 14

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

arXiv:2607. 09767v1 Announce Type: cross Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech.

By Adrien Schneider (M-PSI), Kacper Zabkowski (M-PSI), Anderson Augusma (M-PSI), Fr\'ed\'erique Letu\'e (SAM, SVH), Maria Camila Pinzon (M-PSI), Dominique Vaufreydaz (M-PSI)
arXiv Machine Learning
Sep 24

PHONOS: PHOnetic Neutralization for Online Streaming Applications

PHONOS is a real‑time streaming module for speaker anonymization that neutralizes accent cues by converting non‑native segmental realizations toward a target accent domain. It uses pre‑generated golden utterances that preserve timbre and rhythm, aligning them with silence‑aware DTW and applying zero‑shot voice conversion to supervise a causal accent translator. The system achieves an 81% reduction in non‑native accent confidence, improves accentedness ratings, reduces speaker linkability in embedding space, and operates with ≤241 ms end‑to‑end latency on a single GPU.

By Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna
arXiv AI
Sep 12

Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems

The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.

By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal
arXiv Machine Learning
Sep 3

Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models

The paper introduces a black-box membership inference attack framework tailored for fine-tuned text-to-speech models, addressing challenges in query generation and representation engineering. It evaluates five query types, finding recitation queries most effective, and uses multi-level speech embeddings with temporal alignment for fine-grained comparison. Experiments on CosyVoice2, F5-TTS, and XTTS-v2 trained on VCTK and British Dialect datasets show high privacy leakage, with speaker-level AUC above 0.80 and record-level AUC between 0.80 and 0.90.

By Kunlin Cai, Kaiyuan Zhang, Zihang Xiang, Jinghuai Zhang, Abeer Alwan, Fnu Suya, Yuan Tian
arXiv AI
Sep 4

VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models

The paper introduces VoxPrivacy, a benchmark for assessing interactional privacy in Speech Language Models (SLMs). It evaluates models on a 32‑hour bilingual dataset across three difficulty tiers, revealing that most open‑source SLMs perform near random on conditional privacy decisions and even strong closed‑source systems struggle with proactive privacy inference. The authors also validate these findings on a real‑speech subset and show that fine‑tuning on a 4,000‑hour training set can improve privacy‑preserving capabilities while maintaining robustness.

By Yuxiang Wang, Hongyu Liu, Dekun Chen, Xueyao Zhang, Zhizheng Wu