VISTA (Value-Informed Semantic Trust Arbitration) is a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. It uses a log-odds decomposition to separate emotion expectation from cue diagnosticity, allowing appraisal to change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone, VISTA achieves 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency, and a frozen-backbone probe reaches 0.600 macro CCC for appraisal readout versus 0.505 for emotion-only fine-tuning.
By Jiale Dai, Liuxian Ma, Xiaoke Niu, Wenjing Zhang, Huiying Zhao, Zhaoxiang Liu, Shiguo Lian, Guojie Song
BiCFlow-MER introduces a conditional-flow framework for audio-text multimodal emotion recognition, treating the task as generative evidence transport within a structured emotion space. It disentangles emotion-oriented evidence from speaker style and lexical content, creating a conflict-aware affective condition that guides bidirectional rectified flow to an explicit emotion-space endpoint. The model verifies candidate emotions via adaptive prototype-cloud scoring and backward class-to-condition consistency, achieving superior performance on IEMOCAP, MELD, and the zero-shot CASE benchmark.
By Yanbing Wang, Shenyue Wang, Chunyang Yu
In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary cues from multiple modalities. In audio-text MER, affective cues are often entangled with spea...
arXiv:2608. 04054v1 Announce Type: cross Abstract: Multimodal intent recognition requires understanding not only what textual, acoustic, and visual signals share, but also how they disagree.
By Mohnish Raj, Suraj Kumar, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta
arXiv:2607. 18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance.
By Zilong Huang, Kong Aik Lee, Junjie Li, Zhe Li, Man-Wai Mak
ReH-FUSE is a reliability‑aware hierarchical fusion framework for multimodal emotion recognition in conversation. It uses a decision‑level router to first compare the relative preference between text and audio, then balances this unimodal mixture with a cross‑modal expert, thereby separating unimodal competition from cross‑modal selection. Experiments on IEMOCAP and MELD show that ReH-FUSE achieves state‑of‑the‑art weighted and macro F1 scores, and ablation studies confirm that learned routing outperforms uniform expert averaging and benefits from cross‑modal interaction.
By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen
arXiv:2609.22778v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require exper...
By Yuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn, Zhaonan Wang
arXiv:2608. 04013v1 Announce Type: cross Abstract: Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs.
By Yuntao Shou, Tao Meng, Wei Ai, Keqin Li
arXiv:2609.38182v1 Announce Type: cross
Abstract: Avatar-based multimodal empathetic response generation has emerged as a pivotal capability in human-centric systems, aiming to recognize user emotion...
By Xiaolin Chen, Xuemeng Song, Jinlan Fu, Weili Guan, Mong-Li Lee, Wynne Hsu
arXiv:2608. 03611v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete.
By Chunlei Meng, Jacqueline J. Pang, Pengbin Feng, Zhenyu Yu, Chun Ouyang, Zhongxue Gan
arXiv:2608. 04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer.
By De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma
arXiv:2608. 03475v1 Announce Type: cross Abstract: Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantically conflicting, or disproportionately dominant.
By Suraj Kumar, Mohnish Raj, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta