arXiv:2603.02767v4 Announce Type: replace-cross
Abstract: Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield repre...
By Hanpeng Liu, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, Kaiwen Long, Yaqian Li, Kun He
arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.
By Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu
The paper introduces Inverted Asymmetric Fusion (IAF) to address strong-modality collapse in multimodal learning, where dominant modalities are degraded during fusion. IAF preserves the dominant modality by passing it unchanged and letting weaker modalities attend to it, while also strengthening weaker modalities via Modality-Aware Knowledge Distillation. Experiments on MultiHuSE, UR-FUNNY, and MUStARD show that IAF maintains unimodal performance and improves over the best unimodal baseline by up to 8.25%.
By Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
arXiv:2608. 05000v1 Announce Type: cross Abstract: Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining.
By Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis
arXiv:2609.05916v1 Announce Type: cross
Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...
By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang
PRISM is a training‑free framework that efficiently selects visual instruction data for multimodal large language models by addressing the anisotropy in visual feature distributions, which causes a Global Semantic Drift. By implicitly re‑centering visual semantics, PRISM removes the influence of global background features, cutting data‑selection and model‑tuning time to 30% of conventional pipelines while improving performance across eight multimodal and three language benchmarks, achieving a 101.7% relative gain over baseline models.
By Jinhe Bi, Aniri, Zengjie Jin, Yifan Wang, Danqi Yan, Wenke Huang, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma
arXiv:2604. 00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks.
By Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee
arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.
By Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan
arXiv:2508. 12466v2 Announce Type: replace-cross Abstract: Traditional multimodal learning approaches rely on alignment pre-training to bridge vision and language modalities, typically by projecting visual features into discrete text token spaces using large-scale image--text data.
By Xuhui Zhan, Tyler Derr
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
arXiv:2608.06411v2 Announce Type: replace-cross
Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by...
By Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang, Hao Geng, Minjun Yu
arXiv:2606. 06249v1 Announce Type: cross Abstract: Transformer-based multimodal models rely on attention mechanisms to integrate information across heterogeneous modalities.
By Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello