arXiv Machine Learning By Hebao Zhu, Dongxia Wu

Architecture-Dependent Fusion Pathways in MLLMs

Read the original on arXiv Machine Learning →

The paper investigates how visual and textual information are fused in Multimodal Large Language Models (MLLMs). By analyzing concatenation and native multimodal architectures through alignment decoupling, attention routing, entropy, intrinsic dimensionality, and causal interventions, the authors uncover two distinct fusion pathways: concatenation models use a text‑first, vision‑later strategy, while native models integrate vision and text earlier and reorganize feature spaces. The study also employs visual CKA to test the Platonic Representation Hypothesis, offering a mechanistic view of multimodal fusion and informing architecture‑aware diagnostics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Sep 24

The Alignment Illusion in Multimodal Large Language Models

The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects cross‑modal content integration. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that common scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from intact visual tokens, a phenomenon they term the "alignment illusion." They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task accuracy and reveals when internal geometry diverges from performance.

arXiv AI
Jun 24

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

arXiv:2606. 23885v1 Announce Type: cross Abstract: Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder.

By Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi