arXiv Machine Learning By Diego Torres, Axel Roebel, Nicolas Obin

A Comprehensive Study of Content Representations for Speech Synthesis

Read the original on arXiv Machine Learning →

The paper investigates how different speech content representations—such as SSL features, supervised tokens, posteriorgrams, and neural audio codecs—perform when used to train a generative model that produces audio conditioned only on each representation. By evaluating the generated audio on content, speaker identity, and prosody, the study identifies two regimes: some representations almost fully reconstruct the original audio, while others effectively separate speaker identity. The findings reveal that disentanglement of speaker identity depends on both the training objective and the representation’s information capacity, rather than supervision alone.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 9

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

arXiv:2512. 20978v2 Announce Type: replace-cross Abstract: Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech.

By Haoyang Li, Xuyi Zhuang, Azmat Adnan, Ye Ni, Wei Rao, Shreyas Gopal, Eng Siong Chng, Boon Siew Han, Yuanjin Zheng
arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv AI
3d ago

MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

MeanVoiceFlow2 is a new voice conversion framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. It is trained via conversion distillation from MeanVoiceFlow and real data reconstruction, and further enhanced with diffusion-GAN training, sample mixing, and teacher-guided conditioning augmentation. Experiments on zero-shot voice conversion show that MeanVoiceFlow2 delivers higher perceptual quality and about nine times faster inference than its predecessor while preserving speaker similarity.

By Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo