arXiv:2604. 00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention.
By Junxian Wu, Chenghan Fu, Zhanheng Nie, Daoze Zhang, Bowen Wan, Wanxian Guan, Chuan Yu, Jian Xu, Bo Zheng
arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.
By Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu
arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.
By Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan
arXiv:2508. 16170v2 Announce Type: replace-cross Abstract: MultiModal Recommendation (MMR) systems have emerged as a promising solution for improving recommendation quality by leveraging rich item-side modality information, prompting a surge of diverse methods.
By Xiaoxiong Zhang, Xin Zhou, Zhiwei Zeng, Yongjie Wang, Zhiqi Shen
HMGCLIP is a unified multimodal embedding framework that uses a heterogeneous hypergraph to capture both fine‑grained and coarse‑grained product attributes. By mining structure‑aware hard negatives and aligning multi‑granular semantics at relation and hyperedge levels, it enables a dual‑granularity inference mechanism that dynamically fuses attribute evidence. Experiments on a new fine‑grained e‑commerce dataset and the public MAVE benchmark show that HMGCLIP outperforms strong multimodal encoders, MLLMs, and e‑commerce baselines.
By Qiuyu Zhu, Yi Gao, Zhichao Wan, Mingyang Ma
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
By Fan Xu, Luis A. Leiva
arXiv:2505. 19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities.
By Sanghyuk Chun, Olga Russakovsky
arXiv:2607. 29213v1 Announce Type: cross Abstract: Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent.
By Jiping Liu, Zhongmin Zhang, Zisen Sang, Zhijia Fang, Tao Ouyang, Ma Jiang, Shaopeng Liang, Zeyang Hou, Guodong Cao, Jia Jia
The paper introduces GPUB, a large-scale benchmark for grounded product understanding in e‑commerce livestream videos, featuring 3,000 livestreams, 31K fashion products, and multi‑moment temporal annotations. It defines three evaluation tasks, with the main task (GPrU) requiring simultaneous product identification and moment localization. Existing multimodal models perform poorly on GPrU, prompting the authors to develop UniPro, which improves performance by learning product‑aligned, temporally structured representations.
By Xinyu Zhang, Junjie Chen, Jiawei Ge, Qianlong Li, Libin Ma, Baokun Pan, Yahui Luo
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This...
Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they implicitly encode product information into global embeddings, thereby limiting their ability...
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendatio...