The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.
By Hong-Han Wang, Yuntao Wang, Hu Ding
arXiv:2506.02015v4 Announce Type: replace
Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still strug...
By Yoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi, Sungwoong Kim
The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects cross‑modal content integration. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that common scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from intact visual tokens, a phenomenon they term the "alignment illusion." They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task accuracy and reveals when internal geometry diverges from performance.
arXiv:2606. 31054v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image.
By Zhiyuan Yao, Zheren Fu, Zhixiao Zheng, Jiajun Li, Yi Tu, Zhendong Mao
arXiv:2606. 26387v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text.
By Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, Tianyang Wang, Hao Xu
arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.
By Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan
arXiv:2605. 16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.
By Qinwu Xu
arXiv:2606. 23885v1 Announce Type: cross Abstract: Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder.
By Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi
arXiv:2606. 31876v1 Announce Type: new Abstract: To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space.
By Moreno D'Inc\`a, Massimiliano Mancini, Nicu Sebe
arXiv:2607. 23944v1 Announce Type: new Abstract: Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions.
By Hao Yang, Jin Wang, Xuejie Zhang
arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.
By Xinpeng Dong, Min Zhang, Kairong Han, Xu Tan, Fei Wu, Kun Kuang
The paper introduces Vision-Free Adaptation (VFA), a method that separates multilingual language enhancement from visual alignment in multimodal large language models. VFA fine‑tunes a base LLM on multilingual text to create a multilingual task vector, which is then merged with the vision‑aligned task vector of an existing MLLM. Experiments on five MLLMs and six multilingual benchmarks show consistent gains while preserving multimodal and text‑only performance, and using less than 2% of text data narrows the performance gap to fully multimodal‑trained models.
By Yixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang, Lingjie Jiang, Shaohan Huang, Yun Chen, Guanhua Chen, Furu Wei