arXiv AI

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

arXiv:2606. 14922v1 Announce Type: cross Abstract: For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning.

arXiv AI
Aug 26

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

EmoTra‑TTS introduces a method for smooth intra‑utterance emotion transitions in speech synthesis. It uses a multi‑pass flow blending pipeline, dual‑stage VAD conditioning, and direction‑magnitude decoupled injection to generate frame‑aligned emotional prosody. The system adds only 0.43% more parameters, incurs no latency, and outperforms four state‑of‑the‑art baselines and two commercial systems in emotion transition quality and overall preference tests.

By Tianchi Liu, Zeyang Song, Tianrui Wang, Zhipeng Li, Chenglin Xu, Yiwen Guo
arXiv AI
Aug 24

Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

Poly-InstructTTS is a text‑to‑speech system that learns expressive speech from open‑ended natural‑language instructions using a 1,000‑hour, 1,000‑plus emotion and style annotated audiovisual corpus. The approach employs a prompt‑free GPT with attribute‑based thinking tokens and a flow‑matching module to inject timbre from reference audio, and includes a speaker fine‑tuning procedure to transfer instruction control while preserving speaker persona. Experiments demonstrate strong instruction adherence and expressiveness, with audio demos and an expanded test set available on the project page.

By Junhui Zhang, Qianhui Xu, Qingxiang Guo, Dawei Yang, Ling Miao, Qiangqiang Wang, Yang Song
arXiv Computation and Language
Sep 22

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...

By Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue
arXiv Machine Learning
Aug 17

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

arXiv:2608. 13613v1 Announce Type: cross Abstract: Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions.

By Jiarui Hai, Karan Thakkar, Ke Chen, Yunyun Wang, Jiaqi Su, Rithesh Kumar, Mounya Elhilali, Zeyu Jin
arXiv Computation and Language
Sep 23

Enriching Speech Emotion Representations with Conversational Context

The paper introduces ACERT, a module that incorporates a flexible-length window of conversational context to enhance Speech Emotion Recognition (SER). By capturing emotional evolution across utterances, ACERT outperforms state‑of‑the‑art methods on IEMOCAP, sets a new context‑aware benchmark on SAFE, and achieves strong results on MELD. Ablation studies attribute ACERT’s improvements to emotional and conversational continuity rather than speaker identity or acoustic conditions.

By Arthur Peuvot, Romaric Besan\c{c}on, Ga\"el de Chalendar, Bianca Vieru, Ioana Vasilescu
arXiv Machine Learning
5d ago

A Comprehensive Study of Content Representations for Speech Synthesis

The paper investigates how different speech content representations—such as SSL features, supervised tokens, posteriorgrams, and neural audio codecs—perform when used to train a generative model that produces audio conditioned only on each representation. By evaluating the generated audio on content, speaker identity, and prosody, the study identifies two regimes: some representations almost fully reconstruct the original audio, while others effectively separate speaker identity. The findings reveal that disentanglement of speaker identity depends on both the training objective and the representation’s information capacity, rather than supervision alone.

By Diego Torres, Axel Roebel, Nicolas Obin
arXiv Computation and Language
4d ago

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

arXiv:2606.28249v2 Announce Type: replace-cross Abstract: Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised...

By Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao, Jingyuan Xing, Baiji Liu, Xiangmin Xu