The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv:2607. 05019v1 Announce Type: new Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks.
By Ilya Burenko, Dmitry Vetrov
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
arXiv:2608. 02769v1 Announce Type: cross Abstract: Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance.
By Sagnik Nandy, Samriddha Lahiry, Pragya Sur, Subhabrata Sen
arXiv:2603. 22372v2 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) forecasting.
By Seunghan Lee, Jun Seo, Jaehoon Lee, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, SoonYoung Lee, Wonbin Ahn
arXiv:2505. 19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities.
By Sanghyuk Chun, Olga Russakovsky
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
arXiv:2606. 02659v1 Announce Type: cross Abstract: Multimodal data fusion involves integrating and analyzing information from multiple modalities to uncover latent correlations and complementary patterns, thereby enhancing data processing and decision-making.
By Dong Li, Lingling Zhang, Binghao Han, Linlin Ding, Yue Kou
arXiv:2606. 00959v1 Announce Type: new Abstract: Understanding modality interaction in multimodal large language models (MLLMs) is central to reliable deployment.
By Wanlong Fang, Tianle Zhang, Wen Tao, Alvin Chan
The paper introduces Inverted Asymmetric Fusion (IAF) to address strong-modality collapse in multimodal learning, where dominant modalities are degraded during fusion. IAF preserves the dominant modality by passing it unchanged and letting weaker modalities attend to it, while also strengthening weaker modalities via Modality-Aware Knowledge Distillation. Experiments on MultiHuSE, UR-FUNNY, and MUStARD show that IAF maintains unimodal performance and improves over the best unimodal baseline by up to 8.25%.
By Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
arXiv:2607. 16789v1 Announce Type: new Abstract: Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities.
By Sana Tonekaboni, Viktoria Schuster, Caroline Uhler
arXiv:2606. 11614v1 Announce Type: cross Abstract: Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions.
By Zequn Yang, Yake Wei, Haotian Ni, Zhihao Xu, Di Hu