arXiv:2609.38157v1 Announce Type: cross
Abstract: Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional train...
By Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee, Jinwon Lee, Neha Chachra
arXiv:2606. 07309v1 Announce Type: cross Abstract: Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in a grounded way when the raw audio is already available.
By Iosif Tsangko, Andreas Triantafyllopoulos, Bj\"orn W. Schuller
arXiv:2609.39453v1 Announce Type: cross
Abstract: Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent app...
By Hezhao Zhang, Thomas Hain
arXiv:2606. 27717v1 Announce Type: cross Abstract: Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech.
By Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang, Haonan Chen, Zeyu Jin
arXiv:2608.28932v1 Announce Type: new
Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-o...
By Models Luc Debaupte, Tyler Baumgartner, Brandon Tai, Candice Fan, Bill Wang, Yi Zhong
arXiv:2606.28249v2 Announce Type: replace-cross
Abstract: Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised...
By Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao, Jingyuan Xing, Baiji Liu, Xiangmin Xu
arXiv:2605. 16739v2 Announce Type: replace-cross Abstract: Decoding visual experience from brain activity has advanced substantially, but current brain-to-text systems largely recover semantic content while discarding affect.
By Bilal A. Mohammed, Lin Gu, Ruogu Fang
arXiv:2606. 00129v1 Announce Type: cross Abstract: Large language models (LLMs) have emerged as powerful representation learners whose internal features increasingly align with human cognition.
By Yousef A. Radwan, Xuhui Liu, Kilichbek Haydarov, Yuqian Fu, Mohamed Elhoseiny
The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.
By Tian Fang, Ga\"el Guibon, Davide Buscaldi
The paper introduces an LLM-based framework for continuous dimensional emotion evaluation in multimodal dialogue, combining discrete emotion recognition with Valence-Arousal-Dominance (VAD) assessment on the IEMOCAP dataset. It incorporates acoustic cues as natural language descriptions via the SpeechCueLLM approach and evaluates six models from the LLaMA, GPT, and Qwen families using zero-shot, few-shot, and LoRA fine-tuning. LoRA-fine-tuned LLaMA models outperform prompt-engineered GPT models, achieving a new state-of-the-art Valence CCC of 0.7822, and ablation studies show that textual audio descriptions significantly benefit smaller models.
"whyItMatters":"The study demonstrates that domain adaptation through fine-tuning can surpass larger GPT models in multimodal emotion evaluation, highlighting the importance of tailored training for emotion recognition tasks."
By Yutong Hu, Jinho Choi
arXiv:2512.07571v3 Announce Type: replace
Abstract: This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tune...
By Nicolas Calbucura, Jose Guillen, Valentin Barriere
arXiv:2602. 03420v2 Announce Type: replace-cross Abstract: Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content.
By Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia, Gongping Huang, James Bailey, Ting Dang