← Back to all news
arXiv Computation and Language October 7, 2026 By Seymanur Akti, Alexander Waibel

Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • rag
  • diffusion
  • robotics
  • multimodal

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 19

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

arXiv:2603. 04219v2 Announce Type: replace-cross Abstract: We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis.

By Youngwon Choi, Jinwoo Oh, Hwayeon Kim, Hyeonyu Kim
ragfine-tuningmultimodal
More like this →
arXiv AI
Jun 26

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

arXiv:2606. 26534v1 Announce Type: cross Abstract: Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.

By Tianxin Xie, Chenxing Li, Dong Yu, Li Liu
diffusionreinforcement-learningfine-tuningmultimodalbenchmarks
More like this →
arXiv AI
Jun 16

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

arXiv:2606. 14922v1 Announce Type: cross Abstract: For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning.

By Vinh Dang Quang, Huy Ngo Quang
rag
More like this →
arXiv Machine Learning
Sep 23

Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis

arXiv:2609.25411v1 Announce Type: cross Abstract: Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpola...

By Biel Tura Vecino, Yoach Lacombe, Julian Weber, Zbigniew {\L}atka, Haitong Zhang, Logan Hart, Eren G\"olge
ragmultimodal
More like this →
arXiv Computation and Language
1d ago

Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

arXiv:2610.07647v1 Announce Type: cross Abstract: Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investiga...

By Seymanur Akti, Alexander Waibel
multimodal
More like this →
arXiv AI
Jul 31

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

arXiv:2607. 26742v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.

By Carlos Mu\~noz-Romero, Jose A. Gonzalez-Lopez
diffusionmultimodal
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea