arXiv Machine Learning

CONFER: Conflict-Aware Evidence Negotiation for Regime-Calibrated Weak Supervision in Multimodal Emotion Recognition

arXiv:2608. 07867v1 Announce Type: new Abstract: Multimodal emotion recognition often treats self-reported labels as reliable supervision while overlooking self-report unreliability and cross-modal conflict.

arXiv AI
4d ago

VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict

VISTA (Value-Informed Semantic Trust Arbitration) is a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. It uses a log-odds decomposition to separate emotion expectation from cue diagnosticity, allowing appraisal to change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone, VISTA achieves 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency, and a frozen-backbone probe reaches 0.600 macro CCC for appraisal readout versus 0.505 for emotion-only fine-tuning.

By Jiale Dai, Liuxian Ma, Xiaoke Niu, Wenjing Zhang, Huiying Zhao, Zhaoxiang Liu, Shiguo Lian, Guojie Song
arXiv AI
Sep 24

BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport

BiCFlow-MER introduces a conditional-flow framework for audio-text multimodal emotion recognition, treating the task as generative evidence transport within a structured emotion space. It disentangles emotion-oriented evidence from speaker style and lexical content, creating a conflict-aware affective condition that guides bidirectional rectified flow to an explicit emotion-space endpoint. The model verifies candidate emotions via adaptive prototype-cloud scoring and backward class-to-condition consistency, achieving superior performance on IEMOCAP, MELD, and the zero-shot CASE benchmark.

By Yanbing Wang, Shenyue Wang, Chunyang Yu
arXiv Machine Learning
Sep 15

ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation

ReH-FUSE is a reliability‑aware hierarchical fusion framework for multimodal emotion recognition in conversation. It uses a decision‑level router to first compare the relative preference between text and audio, then balances this unimodal mixture with a cross‑modal expert, thereby separating unimodal competition from cross‑modal selection. Experiments on IEMOCAP and MELD show that ReH-FUSE achieves state‑of‑the‑art weighted and macro F1 scores, and ablation studies confirm that learned routing outperforms uniform expert averaging and benefits from cross‑modal interaction.

By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen