An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis
arXiv:2606. 14922v1 Announce Type: cross Abstract: For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning.
Poly-InstructTTS is a text‑to‑speech system that learns expressive speech from open‑ended natural‑language instructions using a 1,000‑hour, 1,000‑plus emotion and style annotated audiovisual corpus. The approach employs a prompt‑free GPT with attribute‑based thinking tokens and a flow‑matching module to inject timbre from reference audio, and includes a speaker fine‑tuning procedure to transfer instruction control while preserving speaker persona. Experiments demonstrate strong instruction adherence and expressiveness, with audio demos and an expanded test set available on the project page.
arXiv:2606. 14922v1 Announce Type: cross Abstract: For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning.
arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.
arXiv:2510.22588v2 Announce Type: replace-cross Abstract: Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction tha...
arXiv:2502. 16584v2 Announce Type: replace-cross Abstract: Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs).
Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations.
arXiv:2607. 15755v1 Announce Type: cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions.
arXiv:2608.30325v1 Announce Type: new Abstract: Natural-language instructions enable flexible control of synthesized speech, yet emotional TTS systems primarily model a single utterance-level affect,...
arXiv:2608. 10720v1 Announce Type: new Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied.
The paper introduces a scalable method for creating pseudo-triplet data—comprising a reference utterance, a direction text, and a modified utterance—to train direction‑following text‑to‑speech (TTS) systems. It uses an impression‑controllable TTS model to generate style variations and a large language model to generate natural language directions from estimated impression differences. Experiments show that these pseudo‑triplets enable stable speaker‑preserving modifications, and combining them with recorded data further improves direction alignment while maintaining speaker similarity.
arXiv:2608. 02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply.
arXiv:2605.27971v2 Announce Type: replace-cross Abstract: When large language models are fine-tuned to generate persona- or tone-conditioned responses, their output diversity is severely limited--a f...
Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new uttera...