arXiv Machine Learning By Zilong Huang, Kong Aik Lee, Junjie Li, Zhe Li, Man-Wai Mak

EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

Read the original on arXiv Machine Learning →

arXiv:2607. 18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 15

ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation

ReH-FUSE is a reliability‑aware hierarchical fusion framework for multimodal emotion recognition in conversation. It uses a decision‑level router to first compare the relative preference between text and audio, then balances this unimodal mixture with a cross‑modal expert, thereby separating unimodal competition from cross‑modal selection. Experiments on IEMOCAP and MELD show that ReH-FUSE achieves state‑of‑the‑art weighted and macro F1 scores, and ablation studies confirm that learned routing outperforms uniform expert averaging and benefits from cross‑modal interaction.

By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen
arXiv Machine Learning
5d ago

Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis

Reliability-aware Cross-sample Enhancement (RCE) is a framework for multimodal sentiment analysis that tackles noise and missing modalities by first applying an adaptive variational information bottleneck to compress unreliable modality information. It then retrieves high‑confidence, semantically consistent neighbors from a large candidate pool to enrich current representations, and finally fuses cross‑modal interactions through a multilevel reliability‑aware mechanism. Experiments show RCE consistently outperforms state‑of‑the‑art methods in full, noisy, and missing‑modality scenarios.

By Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
arXiv AI
Sep 25

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition proposes a hybrid framework that combines a Transformer and a Graph Attention Network to capture both global semantic information and fine-grained relationships between modalities. The model is evaluated on the IEMOCAP and MELD datasets, achieving weighted F1 scores of 72.45% and 77.37%, respectively, and surpasses state‑of‑the‑art methods. These results suggest that integrating multimodal features with balanced global and local context modeling can provide deeper emotional insights for dialogue emotion recognition.

By Jiaqi Qiao, Yifan Lyu, Xiujuan Xu