arXiv:2608. 03611v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete.
By Chunlei Meng, Jacqueline J. Pang, Pengbin Feng, Zhenyu Yu, Chun Ouyang, Zhongxue Gan
Reliability-aware Cross-sample Enhancement (RCE) is a framework for multimodal sentiment analysis that tackles noise and missing modalities by first applying an adaptive variational information bottleneck to compress unreliable modality information. It then retrieves high‑confidence, semantically consistent neighbors from a large candidate pool to enrich current representations, and finally fuses cross‑modal interactions through a multilevel reliability‑aware mechanism. Experiments show RCE consistently outperforms state‑of‑the‑art methods in full, noisy, and missing‑modality scenarios.
By Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
The paper critiques current optimization-based methods for balancing modalities in Multimodal Sentiment Analysis, arguing they overpromise and underdeliver. It introduces a unified evaluation framework that tests gradient- and loss-based balancing strategies, provides a theoretical diagnosis showing these methods conflate fitting speed with discriminative contribution, and proposes a research agenda for held‑out discriminative modality valuation. Experiments on CMU‑MOSI and CMU‑MOSEI demonstrate that no strategy consistently outperforms Late Concatenation, performance is highly sensitive to hyperparameters, and ratio calibration does not yield reliable gains, highlighting that loss is not utility and gradients are not importance.
By Ioanna Kaffeza, Efthymios Georgiou, Alexandros Potamianos
arXiv:2607. 10599v1 Announce Type: new Abstract: Multimodal sentiment analysis relies on language, visual, and acoustic cues, but utterance-level modality quality may vary due to occlusion, background noise, motion blur, or imperfect transcripts, causing conventional fusion to over-trust unreliable modalities.
By Haoran Ma, Yinfeng Yu, Liejun Wang
arXiv:2607. 08573v1 Announce Type: new Abstract: Multimodal emotion and sentiment recognition is commonly addressed by early fusion, which concatenates modalities before classification, or late fusion, which combines independently trained unimodal predictors.
By Adis Alihodzic, Selma Skopljakovic Hubljar
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
arXiv:2607. 20742v1 Announce Type: new Abstract: Multimodal learning is a robust approach to improve predictive performance in applications such as medical prognosis.
By Mohammad Raahemi, Ali Sekhavati, Alireza Maleki, Hamid Nasiri
arXiv:2608. 03475v1 Announce Type: cross Abstract: Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantically conflicting, or disproportionately dominant.
By Suraj Kumar, Mohnish Raj, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta
ReH-FUSE is a reliability‑aware hierarchical fusion framework for multimodal emotion recognition in conversation. It uses a decision‑level router to first compare the relative preference between text and audio, then balances this unimodal mixture with a cross‑modal expert, thereby separating unimodal competition from cross‑modal selection. Experiments on IEMOCAP and MELD show that ReH-FUSE achieves state‑of‑the‑art weighted and macro F1 scores, and ablation studies confirm that learned routing outperforms uniform expert averaging and benefits from cross‑modal interaction.
By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen
arXiv:2608.29278v1 Announce Type: new
Abstract: Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and read...
By Zhaolu Kang, Meixin Wu, Yu Xue, Yingjie He, Qiming Shi, Lei Wei, Yidi Wang, Richeng Xuan, Zhichao Hu
arXiv:2608. 20019v1 Announce Type: new Abstract: Incomplete multimodal sentiment analysis has garnered significant attention in recent years.
By Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, Zixuan Yu, Wenxi Zhao, Yibei Liu, Qianle Zhang, Yangyang Wu, Mengying Zhu, Meng Xi
arXiv:2607. 27289v1 Announce Type: new Abstract: The promise of multimodal fusion lies in combining complementary sources of evidence, yet more evidence does not always yield a better prediction.
By Yu Chang, Anzhe Cheng, Chenwei Wu, Zhuoran Wang, Jiahao Chen, Tamoghna Chattopadhyay, Sophia I. Thomopoulos, Paul M. Thompson, Liyue Shen, Paul Bogdan