arXiv AI By Adrien Schneider (M-PSI), Kacper Zabkowski (M-PSI), Anderson Augusma (M-PSI), Fr\'ed\'erique Letu\'e (SAM, SVH), Maria Camila Pinzon (M-PSI), Dominique Vaufreydaz (M-PSI)

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

Read the original on arXiv AI →

arXiv:2607. 09767v1 Announce Type: cross Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Anonymization, Not Elimination: Utility-Preserved Speech Anonymization

The paper introduces a two‑stage speech anonymization framework that preserves both linguistic content and acoustic identity. It replaces personally identifiable information using a generative editing model and applies a flow‑matching anonymization technique (F3‑VA) to create diverse, distinct anonymized speakers. The authors evaluate privacy with speaker verification metrics and utility by training ASR, TTS, and SER models from scratch, showing stronger privacy protection with minimal utility loss compared to existing baselines.

By Yunchong Xiao, Yuxiang Zhao, Ziyang Ma, Shuai Wang, Kai Yu, Jiachun Liao, Xie Chen
arXiv Computation and Language
Aug 28

Your Voice Cloning System is Secretly a Voice Anonymizer

The paper demonstrates that the multilingual voice cloning model XTTSv2 can be repurposed for speaker anonymization without retraining. By conditioning on a pseudo-speaker and using an iterative refinement strategy, the authors balance privacy and intelligibility, achieving near‑optimal privacy (EER ≈ 0.49) and competitive speech quality across seven European languages. The method outperforms dedicated anonymization baselines and requires no language‑specific training.

By Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu
arXiv Computation and Language
Sep 4

SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).

By Eunseo Choi, Hyunku Kang, Chanwoo Kim
arXiv Machine Learning
Sep 24

PHONOS: PHOnetic Neutralization for Online Streaming Applications

PHONOS is a real‑time streaming module for speaker anonymization that neutralizes accent cues by converting non‑native segmental realizations toward a target accent domain. It uses pre‑generated golden utterances that preserve timbre and rhythm, aligning them with silence‑aware DTW and applying zero‑shot voice conversion to supervise a causal accent translator. The system achieves an 81% reduction in non‑native accent confidence, improves accentedness ratings, reduces speaker linkability in embedding space, and operates with ≤241 ms end‑to‑end latency on a single GPU.

By Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna