arXiv:2607. 18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance.
By Zilong Huang, Kong Aik Lee, Junjie Li, Zhe Li, Man-Wai Mak
The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
By Xu Lin, Ke Wang, Hui Kang, Xinying Wang
arXiv:2608.20905v1 Announce Type: new
Abstract: Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue...
By Yi Zheng, Yifan Xu, Yan Zhou, Hejia Chen, Chunyu Qiang, Xiaoqiang Liu, Xiaohan Li, Shenze Huang, Yue Zhang, Guoying Zhao, Pengfei Wan
ReH-FUSE is a reliability‑aware hierarchical fusion framework for multimodal emotion recognition in conversation. It uses a decision‑level router to first compare the relative preference between text and audio, then balances this unimodal mixture with a cross‑modal expert, thereby separating unimodal competition from cross‑modal selection. Experiments on IEMOCAP and MELD show that ReH-FUSE achieves state‑of‑the‑art weighted and macro F1 scores, and ablation studies confirm that learned routing outperforms uniform expert averaging and benefits from cross‑modal interaction.
By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen
The paper introduces Emo-DVS, a large-scale, multimodal dataset combining event camera, audio, and text data for emotion recognition, designed to mitigate privacy concerns associated with RGB cameras. It proposes the Information‑Guided Gated Fusion (IGF) framework, which pre‑trains an event encoder on the dataset’s FAU subset, adaptively gates modalities to reduce noise, and aligns cross‑modal representations via mutual information maximization. Experiments show that IGF outperforms existing methods on this challenging tri‑modal benchmark.
By Jiaqi Chen, Qinfu Xu, Hao Zhuang, Liyuan Pan
DiaRelay introduces a lightweight adapter that lets large language models maintain a constant‑size dialogue‑level memory for emotion recognition in conversation. It builds on LoRA by adding a Selective Relay Memory Transition that aggregates useful historical evidence into a bounded memory, and a Dual‑axis Relay Memory Read that uses this memory to modulate low‑rank feature transformations. Experiments show DiaRelay achieves state‑of‑the‑art weighted F1 and accuracy on MELD with only 7.1 M additional trainable parameters, while also performing competitively on IEMOCAP.
By Zihao Zhou, Bin Yang, Jinghui Qin, Kebing Jin
The paper introduces EmoVec, a lightweight framework that enables controllable affective generation in large language models by steering latent vectors. EmoVec identifies emotion-specific directions from paired neutral and emotion-conditioned responses using contrastive activation addition, then refines these directions through task-specific debiasing and principal subspace removal. During inference, the refined vectors are injected into the final residual stream with static or scenario-adaptive scaling, allowing continuous control over emotional intensity without updating model weights, and experiments across three LLMs and eight emotions demonstrate improved emotional salience while preserving semantic content, fluency, and coherence.
By Xixian Yong, Siyuan Chang, Yingying Zhang, Xian Wu, Xiao Zhou
arXiv:2607. 12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition.
By Aleksei Bakin, Andrey V. Savchenko
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video.
arXiv:2609.06188v1 Announce Type: new
Abstract: Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic...
By Pengfei Shao, Jisheng Dang, Jiawen Fang, Ning Liu, Wencan Zhang, Bimei Wang, Jingwen Zhao, Jianhuang Lai, Qi Tian, Tat-Seng Chua
arXiv:2607. 14683v1 Announce Type: new Abstract: Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction.
By Hao Yang, Yanyan Zhao, Kewei Zhao, Hongbo Zhang, Tian Zheng, Yusheng Liu, Xing Fu, Bichen Wang, Yu Zhang, Hao He, Zhen Wu, Xuda Zhi, Yongbo Huang, Bing Qin
arXiv:2606. 00851v1 Announce Type: cross Abstract: Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues.
By Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani