arXiv Computation and Language By Yinfeng Wang, Zhiyuan Yao, Zheren Fu, Lei Zhang, Zhendong Mao

When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

Read the original on arXiv Computation and Language →

arXiv:2608. 19208v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.