BiCFlow-MER introduces a conditional-flow framework for audio-text multimodal emotion recognition, treating the task as generative evidence transport within a structured emotion space. It disentangles emotion-oriented evidence from speaker style and lexical content, creating a conflict-aware affective condition that guides bidirectional rectified flow to an explicit emotion-space endpoint. The model verifies candidate emotions via adaptive prototype-cloud scoring and backward class-to-condition consistency, achieving superior performance on IEMOCAP, MELD, and the zero-shot CASE benchmark.
By Yanbing Wang, Shenyue Wang, Chunyang Yu
The paper introduces Affect-Prototype Guided Fusion (APCF), a framework for open‑vocabulary multimodal emotion recognition that handles incomplete and unsynchronized modal data. APCF builds an affect‑prototype library to model how different emotions contribute across modalities, enabling dynamic fusion of available features. The fused representations are then decoded by an LLM to generate open‑vocabulary emotion labels, achieving superior performance on OV‑MERD+ and MER‑FG datasets compared to existing methods.
By Yichi Zhang, Shenyue Wang, Jing Luo, Chunyang Yu, Xinyu Yang
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
arXiv:2609.06188v1 Announce Type: new
Abstract: Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic...
By Pengfei Shao, Jisheng Dang, Jiawen Fang, Ning Liu, Wencan Zhang, Bimei Wang, Jingwen Zhao, Jianhuang Lai, Qi Tian, Tat-Seng Chua
The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
By Xu Lin, Ke Wang, Hui Kang, Xinying Wang
arXiv:2608. 04013v1 Announce Type: cross Abstract: Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs.
By Yuntao Shou, Tao Meng, Wei Ai, Keqin Li
arXiv:2601. 07565v2 Announce Type: replace-cross Abstract: Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, while simultaneously addressing discrete emotion recognition and continuous sentiment analysis.
By Jiaqi Qiao, Xinran Li, Yifan Lyu, Xiujuan Xu, Liu Yu
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations.
arXiv:2607. 18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance.
By Zilong Huang, Kong Aik Lee, Junjie Li, Zhe Li, Man-Wai Mak
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.
arXiv:2608.30726v1 Announce Type: new
Abstract: Multimodal Sentiment Analysis (MSA) is a fundamental component of affective computing that aims to decipher complex emotional states by integrating ver...
By Xiaode Chen, Jiakang Yu, Hongtao Deng, Huina Qu, Xun Zhu, Yinxia Lou
arXiv:2607. 12787v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc.
By Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M. Jose, Xuri Ge