arXiv AI By Yang Qiao, Yuntong Hu, Bowen Zhu, Hasibul Haque, Liang Zhao

Multimodal Representation Learning Conditioned on Semantic Relations

Read the original on arXiv AI →

Multimodal representation learning has largely relied on contrastive models like CLIP that produce a single embedding per sample, which limits their ability to capture relation-dependent relevance. The proposed Relation-Conditioned Multimodal Learning (RCML) framework explicitly conditions embeddings on natural‑language relation descriptions, enabling the same sample to be represented differently under various relational contexts. RCML builds relation‑aware training pairs, incorporates a relation‑conditioned module, and uses a unified contrastive objective to jointly model cross‑modal alignment and relation‑induced structure, achieving superior performance on retrieval and classification tasks across zero‑shot, fine‑tuned, and out‑of‑domain settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Generalized Multimodal Foundation Model

The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.

By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv AI
Sep 18

Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning

The paper introduces Lens, a training‑free framework that aligns multimodal representations with the semantic perspective required by downstream tasks. Lens uses a task‑specific readout phrase to anchor the perspective and then aggregates token states after the full input, ensuring the extracted representation reflects task‑conditioned evidence integration rather than generic salient content. The method achieves a Precision@1 of 63.9 across 36 MMEB datasets, outperforming the nearest training‑free baseline by 10.2 points.

By Xinran Liu, Shouqian Shi, Yixian Chen, Ruizhi Chen, Xin-Wei Yao, Sheng Zhong
arXiv Machine Learning
Jun 3

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

arXiv:2603. 01471v3 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.

By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi