arXiv Machine Learning By Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai

Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis

Read the original on arXiv Machine Learning →

Reliability-aware Cross-sample Enhancement (RCE) is a framework for multimodal sentiment analysis that tackles noise and missing modalities by first applying an adaptive variational information bottleneck to compress unreliable modality information. It then retrieves high‑confidence, semantically consistent neighbors from a large candidate pool to enrich current representations, and finally fuses cross‑modal interactions through a multilevel reliability‑aware mechanism. Experiments show RCE consistently outperforms state‑of‑the‑art methods in full, noisy, and missing‑modality scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction

The paper presents a method for robust multimodal sentiment analysis that handles incomplete or noisy modalities. It introduces a completeness estimation technique to measure how much sentiment-relevant information remains in partial data, guiding the reconstruction of missing semantics. A joint training strategy stabilizes multi-task learning for sentiment prediction and completeness estimation, and experiments on three benchmark datasets show improved semantic reconstruction and sentiment accuracy.

By Han-Jun Choi, Byunggill Joe, Saim Shin, Jin Yea Jang
arXiv Computation and Language
6d ago

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

SemMSA introduces a latent semantic‑aided framework for multimodal sentiment analysis that leverages large language models to generate rich sentiment‑relevant semantics. The method employs Cross‑modal Semantic Refinement (CSR) to fuse visual, acoustic, and language features in a frozen LLM embedding space, and Cross‑modal Spectral Alignment (CSA) to align these refined semantics with all modalities via spectral enhancement of kernel Gram matrices. Experiments on SIMS, MOSI, and MOSEI benchmarks show that SemMSA achieves state‑of‑the‑art performance.

By Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang
arXiv AI
Jul 14

MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis

arXiv:2607. 10599v1 Announce Type: new Abstract: Multimodal sentiment analysis relies on language, visual, and acoustic cues, but utterance-level modality quality may vary due to occlusion, background noise, motion blur, or imperfect transcripts, causing conventional fusion to over-trust unreliable modalities.

By Haoran Ma, Yinfeng Yu, Liejun Wang