arXiv AI

Information-Theoretic Decomposition for Multimodal Interaction Learning

arXiv:2606. 11614v1 Announce Type: cross Abstract: Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions.

arXiv Machine Learning
Sep 22

Generalized Multimodal Foundation Model

The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.

By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv Machine Learning
Jun 10

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.

By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang
arXiv Computation and Language
3d ago

Fusion Anything: A Generalized Multimodal Foundation Model

The paper introduces Fusion Anything Model (FAM), a foundation model designed for generalized multimodal data fusion that can handle arbitrary modality combinations and prediction tasks. FAM is trained on large-scale synthetic multimodal datasets generated via Structural Multimodal Causal Models (SMCMs), enabling it to encode transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show that FAM performs competitively with specialized models without requiring task-specific adaptation.

By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv Computer Vision
3d ago

Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback

The paper introduces MEQ, a mutual feedback architecture that iteratively refines two multimodal inputs into coupled embeddings, each embedding incorporating information from the other. By continuously exchanging information between the modalities, the model converges to a fixed point that improves representation quality. Experiments on classification and visual grounding tasks show that MEQ achieves competitive or superior performance compared to concatenation-based baselines, and qualitatively enhances visual grounding when paired with complementary modalities.

By Ho-min Park, Byungkon Kang
arXiv Computer Vision
2d ago

Gestalt: Large Multimodal Interplay Model

arXiv:2610.00576v1 Announce Type: new Abstract: In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal...

By Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, Ji-Rong Wen, Yake Wei, Di Hu