arXiv:2607. 27660v1 Announce Type: new Abstract: Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning.
By Rishabh Iyer, Truong Pham, Anay Majee
arXiv:2609.23533v1 Announce Type: new
Abstract: Multimodal classifiers can converge to modality-dominant solutions in which one modality dominates the joint prediction, suppressing the learning of ot...
By Zechang Xiong, Da Li, Rong Yin, Kexin Tang, Biao Yang, Pengyuan Li, Wenkang Kong, Yulan Hu, Shengyu Zhu, Hao Peng
Multimodal representation learning seeks shared representations for cross-modal retrieval and knowledge transfer. Hub-based binding reduces pairwise supervision costs, but separate hub connections can...
arXiv:2608. 10857v1 Announce Type: new Abstract: Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning.
By Viktoria Schuster, Sana Tonekaboni, Caroline Uhler
arXiv:2609.17094v1 Announce Type: new
Abstract: Multimodal representation learning seeks shared representations for cross-modal retrieval and knowledge transfer. Hub-based binding reduces pairwise su...
By Ying Guo, Haidong Chen, Linrui Xu, Xiaohao Liu, Chuancheng Shi, Canran Xiao, Dan Zhang, Fei Shen, Li Shen, Tat-Seng Chua
arXiv:2606. 04280v1 Announce Type: cross Abstract: Contrastive learning has become a leading paradigm for self-supervised representation learning, yet the conditions under which it recovers meaningful latent geometry remain incompletely understood.
By Justinas Zaliaduonis, Patrick Putzky, Till Richter, Sergios Gatidis
The paper introduces MEQ, a mutual feedback architecture that iteratively refines two multimodal inputs into coupled embeddings, each embedding incorporating information from the other. By continuously exchanging information between the modalities, the model converges to a fixed point that improves representation quality. Experiments on classification and visual grounding tasks show that MEQ achieves competitive or superior performance compared to concatenation-based baselines, and qualitatively enhances visual grounding when paired with complementary modalities.
By Ho-min Park, Byungkon Kang
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
By Fan Xu, Luis A. Leiva
arXiv:2608.21443v1 Announce Type: new
Abstract: Estimating interpretable conditional-dependence structures from multimodal visual-linguistic features remains largely unexplored. We propose CM-GLasso...
By Fei Wang, Yutong Zhang, Yang Ye, Jinxian Chen, Wang Wenshuai, Xiong Wang
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.
The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.
By Aditya Sharma, Divya Saxena
arXiv:2606. 16639v1 Announce Type: new Abstract: Multimodal learning exploits complementary information across heterogeneous modalities.
By Ankush Pratap Singh, Houwei Cao, Yong Liu