The paper proposes a Mixture-of-Bottleneck (MoB) framework for video-based multimodal sentiment analysis that treats sentiment as an ordinal regression problem, splitting it into polarity recognition and intensity prediction. MoB assigns modality‑specific latent experts to each sub‑task, learns compact, task‑relevant representations via an information bottleneck, and fuses these experts with a multimodal bottleneck routing module and hard mining strategy. Experiments on four datasets and language models demonstrate that MoB captures fine‑grained intra‑ and inter‑modal dynamics, improving performance and enabling more trustworthy localization of nuanced sentiment signals.
By Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan
arXiv:2608. 20019v1 Announce Type: new Abstract: Incomplete multimodal sentiment analysis has garnered significant attention in recent years.
By Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, Zixuan Yu, Wenxi Zhao, Yibei Liu, Qianle Zhang, Yangyang Wu, Mengying Zhu, Meng Xi
arXiv:2606. 01323v1 Announce Type: cross Abstract: Aspect-Based Sentiment Analysis (ABSA) encompasses seven distinct subtasks, each focusing on different extracted elements.
By Shu Long, Yanglei Gan, Xuchuan Zhou
SemMSA introduces a latent semantic‑aided framework for multimodal sentiment analysis that leverages large language models to generate rich sentiment‑relevant semantics. The method employs Cross‑modal Semantic Refinement (CSR) to fuse visual, acoustic, and language features in a frozen LLM embedding space, and Cross‑modal Spectral Alignment (CSA) to align these refined semantics with all modalities via spectral enhancement of kernel Gram matrices. Experiments on SIMS, MOSI, and MOSEI benchmarks show that SemMSA achieves state‑of‑the‑art performance.
By Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang
The paper presents a method for robust multimodal sentiment analysis that handles incomplete or noisy modalities. It introduces a completeness estimation technique to measure how much sentiment-relevant information remains in partial data, guiding the reconstruction of missing semantics. A joint training strategy stabilizes multi-task learning for sentiment prediction and completeness estimation, and experiments on three benchmark datasets show improved semantic reconstruction and sentiment accuracy.
By Han-Jun Choi, Byunggill Joe, Saim Shin, Jin Yea Jang
arXiv:2606. 15694v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content.
By Hangling Xie
arXiv:2608. 16201v1 Announce Type: new Abstract: Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision.
By Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao
arXiv:2608.30425v1 Announce Type: new
Abstract: Cross-lingual aspect-based sentiment analysis (ABSA) transfers knowledge from a source language with annotated data to a target language, enabling fine...
By Jakub \v{S}m\'{i}d, Pavel P\v{r}ib\'{a}\v{n}, Pavel Kr\'{a}l
arXiv:2601. 07565v2 Announce Type: replace-cross Abstract: Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, while simultaneously addressing discrete emotion recognition and continuous sentiment analysis.
By Jiaqi Qiao, Xinran Li, Yifan Lyu, Xiujuan Xu, Liu Yu
arXiv:2608. 19971v1 Announce Type: new Abstract: Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues.
By Zhifa Geng, Subin Huang, Hao Guo, Junjie Chen, Sanmin Liu, Chao Kong
Reliability-aware Cross-sample Enhancement (RCE) is a framework for multimodal sentiment analysis that tackles noise and missing modalities by first applying an adaptive variational information bottleneck to compress unreliable modality information. It then retrieves high‑confidence, semantically consistent neighbors from a large candidate pool to enrich current representations, and finally fuses cross‑modal interactions through a multilevel reliability‑aware mechanism. Experiments show RCE consistently outperforms state‑of‑the‑art methods in full, noisy, and missing‑modality scenarios.
By Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision. While large language models (LLMs) offer strong semantic...