arXiv:2603. 01471v3 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.
By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi
arXiv:2609.37225v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods eith...
By Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong
CausalEmbed is an auto‑regressive method for generating compact multi‑vector embeddings in visual document retrieval. By using iterative margin loss during contrastive training, it reduces the number of visual tokens needed by 30‑155× while keeping performance competitive across different backbones and benchmarks. The approach offers efficient training, scalable test‑time performance, and a flexible scaling strategy for multi‑vector representations.
By Jiahao Huo, Yu Huang, Yibo Yan, Ye Pan, Kening Zheng, Wei-Chieh Huang, Yi Cao, Mingdong Ou, Philip S. Yu, Xuming Hu
arXiv:2604.22280v4 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that...
By Peixi Wu, Ke Mei, Feipeng Ma, Bosong Chai, Zhibin Lan, Chenxi Zhao, Shannan Yan, Jie Chen, Zhangchi Hu, Yansong Peng, Bo Lin, Junjie Zhou, Dacheng Yin, Tianyi Wang, Fengyun Rao, Jing Lyu, Hebei Li, Xiaoyan Sun
arXiv:2607. 16305v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding.
By Zeyu Xu, Xingzhong Hou, Pengkai Guo, Siling Lin, Xiao Xu, Menghua Zhai, Haoyu Chen, Yunke Zhang, Fei Huang
The paper introduces Lens, a training‑free framework that aligns multimodal representations with the semantic perspective required by downstream tasks. Lens uses a task‑specific readout phrase to anchor the perspective and then aggregates token states after the full input, ensuring the extracted representation reflects task‑conditioned evidence integration rather than generic salient content. The method achieves a Precision@1 of 63.9 across 36 MMEB datasets, outperforming the nearest training‑free baseline by 10.2 points.
By Xinran Liu, Shouqian Shi, Yixian Chen, Ruizhi Chen, Xin-Wei Yao, Sheng Zhong