arXiv Computation and Language

Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

arXiv Machine Learning
Sep 29

A Comprehensive Study of Content Representations for Speech Synthesis

The paper investigates how different speech content representations—such as SSL features, supervised tokens, posteriorgrams, and neural audio codecs—perform when used to train a generative model that produces audio conditioned only on each representation. By evaluating the generated audio on content, speaker identity, and prosody, the study identifies two regimes: some representations almost fully reconstruct the original audio, while others effectively separate speaker identity. The findings reveal that disentanglement of speaker identity depends on both the training objective and the representation’s information capacity, rather than supervision alone.

By Diego Torres, Axel Roebel, Nicolas Obin
Hugging Face Trending Papers
Jun 4

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

We propose GLASS, a framework for composable acoustic style control in zero-shot autoregressive text-to-speech (TTS) that learns controls from post-generation rewards rather than style labels. In zero-shot TTS, a speaker prompt often entangles speaker identity with prosodic attributes such as speaking rate and pitch, making it difficult to change style without changing the prompt itself.