arXiv AI By Carlos Mu\~noz-Romero, Jose A. Gonzalez-Lopez

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

Read the original on arXiv AI →

arXiv:2607. 26742v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.