arXiv Machine Learning

From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

arXiv:2606. 08843v1 Announce Type: cross Abstract: We present a voice conversion (VC) framework that utilizes K-Nearest Neighbors (KNN) retrieval over WavLM representations to align non-parallel source and target speech, constructing synthetic training pairs for supervised learning.

arXiv AI
Sep 7

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

X-VC is a zero‑shot streaming voice conversion system that performs one‑step conversion directly in the latent space of a pretrained neural codec. It employs a dual‑conditioning acoustic converter that jointly models source codec latents and target acoustic conditions, while using adaptive normalization to inject utterance‑level speaker information. The model is trained with generated paired data and a role‑assignment strategy, and uses a chunkwise inference scheme with overlap smoothing to achieve low‑latency streaming inference, achieving superior WER, speaker similarity, and real‑time factor on the Seed‑TTS‑Eval benchmark.

By Qixi Zheng, Yuxiang Zhao, Tianrui Wang, Wenxi Chen, Kele Xu, Yikang Li, Qinyuan Cheng, Xipeng Qiu, Kai Yu, Xie Chen
arXiv Machine Learning
Sep 24

PHONOS: PHOnetic Neutralization for Online Streaming Applications

PHONOS is a real‑time streaming module for speaker anonymization that neutralizes accent cues by converting non‑native segmental realizations toward a target accent domain. It uses pre‑generated golden utterances that preserve timbre and rhythm, aligning them with silence‑aware DTW and applying zero‑shot voice conversion to supervise a causal accent translator. The system achieves an 81% reduction in non‑native accent confidence, improves accentedness ratings, reduces speaker linkability in embedding space, and operates with ≤241 ms end‑to‑end latency on a single GPU.

By Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna
arXiv Computation and Language
Sep 16

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

The paper proposes a method called language orthogonalization to improve zero‑shot cross‑lingual audio deepfake detection. By removing language‑dependent variation from self‑supervised speech models using a target‑free ridge map on language‑identification embeddings, the approach consistently lowers equal error rates across six languages and six model backbones. The gains are larger when the target language is more distant in the language‑identification space.

By Minu Kim, Ji Sub Um, Hoirin Kim
arXiv AI
Aug 18

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.

By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov