arXiv:2606. 07293v1 Announce Type: cross Abstract: Speech Emotion Conversion (SEC) aims to transform the emotion of a source utterance into a target emotion while preserving content and speaker identity.
By Constantin Alexander Auga
EmoTra‑TTS introduces a method for smooth intra‑utterance emotion transitions in speech synthesis. It uses a multi‑pass flow blending pipeline, dual‑stage VAD conditioning, and direction‑magnitude decoupled injection to generate frame‑aligned emotional prosody. The system adds only 0.43% more parameters, incurs no latency, and outperforms four state‑of‑the‑art baselines and two commercial systems in emotion transition quality and overall preference tests.
By Tianchi Liu, Zeyang Song, Tianrui Wang, Zhipeng Li, Chenglin Xu, Yiwen Guo
arXiv:2602. 03420v2 Announce Type: replace-cross Abstract: Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content.
By Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia, Gongping Huang, James Bailey, Ting Dang
arXiv:2607. 15755v1 Announce Type: cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions.
By Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu, Haizhou Li
arXiv:2606. 09837v1 Announce Type: cross Abstract: Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis.
By Yue Zhao, Hongyan Li, Yong Chen, Luo Ji
Poly-InstructTTS is a text‑to‑speech system that learns expressive speech from open‑ended natural‑language instructions using a 1,000‑hour, 1,000‑plus emotion and style annotated audiovisual corpus. The approach employs a prompt‑free GPT with attribute‑based thinking tokens and a flow‑matching module to inject timbre from reference audio, and includes a speaker fine‑tuning procedure to transfer instruction control while preserving speaker persona. Experiments demonstrate strong instruction adherence and expressiveness, with audio demos and an expanded test set available on the project page.
By Junhui Zhang, Qianhui Xu, Qingxiang Guo, Dawei Yang, Ling Miao, Qiangqiang Wang, Yang Song