arXiv Machine Learning

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

arXiv:2607. 00946v1 Announce Type: cross Abstract: While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood.

arXiv AI
Aug 26

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

EmoTra‑TTS introduces a method for smooth intra‑utterance emotion transitions in speech synthesis. It uses a multi‑pass flow blending pipeline, dual‑stage VAD conditioning, and direction‑magnitude decoupled injection to generate frame‑aligned emotional prosody. The system adds only 0.43% more parameters, incurs no latency, and outperforms four state‑of‑the‑art baselines and two commercial systems in emotion transition quality and overall preference tests.

By Tianchi Liu, Zeyang Song, Tianrui Wang, Zhipeng Li, Chenglin Xu, Yiwen Guo
arXiv Computation and Language
Aug 27

Controllable Affective Generation via Latent Vector Steering

The paper introduces EmoVec, a lightweight framework that enables controllable affective generation in large language models by steering latent vectors. EmoVec identifies emotion-specific directions from paired neutral and emotion-conditioned responses using contrastive activation addition, then refines these directions through task-specific debiasing and principal subspace removal. During inference, the refined vectors are injected into the final residual stream with static or scenario-adaptive scaling, allowing continuous control over emotional intensity without updating model weights, and experiments across three LLMs and eight emotions demonstrate improved emotional salience while preserving semantic content, fluency, and coherence.

By Xixian Yong, Siyuan Chang, Yingying Zhang, Xian Wu, Xiao Zhou
arXiv Machine Learning
Sep 16

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
arXiv AI
Sep 18

Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

The paper introduces a discriminative adaptation for SpeechLLMs that reads the hidden state of the final prompt token via a simple classification head, enabling emotion recognition in a single forward pass without altering the backbone. This approach replaces the generative decoder, which can produce out‑of‑set labels and favor frequent classes, with a controlled comparison between generative and discriminative inference. Experiments on IEMOCAP show improved Macro F1 scores, elimination of hallucinations, and larger gains on realistic ASR transcripts, while revealing that emotion directions encode indirect associations reflecting web‑scale text biases.

By Hasindri Watawana, Sergio Burdisso, Esa\'u Villatoro-Tello, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
arXiv Machine Learning
5d ago

A Comprehensive Study of Content Representations for Speech Synthesis

The paper investigates how different speech content representations—such as SSL features, supervised tokens, posteriorgrams, and neural audio codecs—perform when used to train a generative model that produces audio conditioned only on each representation. By evaluating the generated audio on content, speaker identity, and prosody, the study identifies two regimes: some representations almost fully reconstruct the original audio, while others effectively separate speaker identity. The findings reveal that disentanglement of speaker identity depends on both the training objective and the representation’s information capacity, rather than supervision alone.

By Diego Torres, Axel Roebel, Nicolas Obin
arXiv AI
Sep 18

Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis

The paper introduces Live-ProsodyJudge (LPJ), a cost‑effective pairwise evaluator distilled from Gemini for assessing fine‑grained prosody in live streaming TTS. It identifies a flaw called verdict coupling, where multi‑dimensional scores collapse into a single preference, and proposes Decoupled‑Live‑ProsodyJudge (D‑LPJ) to eliminate this issue through masking and a span‑local GRPO strategy. Experiments show LPJ outperforms a single Gemini call in accuracy, and D‑LPJ provides independent dimension judgments, achieving high alignment with human top‑3 selections in a TTS candidate tournament.

By Zifan Guan, Longyu Lu, Junan Zhang, Zhizheng Wu, Meiguang Jin, Junfeng Ma