arXiv Machine Learning By Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

Read the original on arXiv Machine Learning →

arXiv:2608. 16201v1 Announce Type: new Abstract: Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 25

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

SemMSA introduces a latent semantic‑aided framework for multimodal sentiment analysis that leverages large language models to generate rich sentiment‑relevant semantics. The method employs Cross‑modal Semantic Refinement (CSR) to fuse visual, acoustic, and language features in a frozen LLM embedding space, and Cross‑modal Spectral Alignment (CSA) to align these refined semantics with all modalities via spectral enhancement of kernel Gram matrices. Experiments on SIMS, MOSI, and MOSEI benchmarks show that SemMSA achieves state‑of‑the‑art performance.

By Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang
arXiv Machine Learning
5d ago

Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis

Reliability-aware Cross-sample Enhancement (RCE) is a framework for multimodal sentiment analysis that tackles noise and missing modalities by first applying an adaptive variational information bottleneck to compress unreliable modality information. It then retrieves high‑confidence, semantically consistent neighbors from a large candidate pool to enrich current representations, and finally fuses cross‑modal interactions through a multilevel reliability‑aware mechanism. Experiments show RCE consistently outperforms state‑of‑the‑art methods in full, noisy, and missing‑modality scenarios.

By Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
arXiv Computation and Language
Sep 17

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

The paper proposes a Mixture-of-Bottleneck (MoB) framework for video-based multimodal sentiment analysis that treats sentiment as an ordinal regression problem, splitting it into polarity recognition and intensity prediction. MoB assigns modality‑specific latent experts to each sub‑task, learns compact, task‑relevant representations via an information bottleneck, and fuses these experts with a multimodal bottleneck routing module and hard mining strategy. Experiments on four datasets and language models demonstrate that MoB captures fine‑grained intra‑ and inter‑modal dynamics, improving performance and enabling more trustworthy localization of nuanced sentiment signals.

By Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan