arXiv Machine Learning By Andrew Stuart, Florian Wolf

Expressivity In Multimodal Contrastive Learning

Read the original on arXiv Machine Learning →

The paper investigates the expressive power of multimodal contrastive learning architectures by treating them as parameterized families of joint density estimators. It shows that the classic two‑tower CLIP model is a universal approximator for two modalities, while a common extension that sums pairwise similarities fails to approximate arbitrary joint distributions when three or more modalities are involved, though it can match all pairwise conditionals. To address this limitation, the authors introduce Hadamard‑CLIP, which adds a single learned weight vector to restore universal approximation for any number of modalities while retaining CLIP’s efficient retrieval capabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 8

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.

By Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan