arXiv:2603.04419v3 Announce Type: replace-cross
Abstract: Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify...
By Murad Farzulla
The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.
By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv:2607. 20092v1 Announce Type: cross Abstract: Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful.
By Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani
arXiv:2608. 19208v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored.
By Yinfeng Wang, Zhiyuan Yao, Zheren Fu, Lei Zhang, Zhendong Mao
arXiv:2609.00293v1 Announce Type: new
Abstract: We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in contex...
By Athulith Paraselli, Etha Tianze Hua, Ellie Pavlick
arXiv:2608. 00410v2 Announce Type: replace Abstract: Human language is highly polysemous.
By Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths
arXiv:2609.07474v2 Announce Type: replace
Abstract: Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should...
By Peng Xie, Amr Alanwar
The paper introduces a perception interface that separates vision from language in vision‑language models. A frozen perception stack detects objects, a deterministic semantic serializer converts the perceived state into text, and a standard text‑only large language model (LLM) answers questions. Experiments on a campus‑robot benchmark show that this serialized interface outperforms a zero‑shot VLM of the same language‑model size, especially as the language model shrinks, and that the advantage persists under paraphrase and different supervision regimes.
By Cong Xu, Ravi Sankar
arXiv:2509. 22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect.
By Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
The paper introduces Visual Retrieval Heads (VRHs), a small fraction of attention heads in vision‑language models that are causally responsible for grounding text descriptions to image regions. By recasting head‑scoring methods and evaluating across eleven VLMs and five benchmarks, the authors show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect. VRHs generalize across various visual reference tasks, preserve output format while corrupting localization, and transfer causally across models sharing an LLM backbone.
By Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung
arXiv:2606. 17389v1 Announce Type: cross Abstract: Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical.
By Logan Mann, Yi Xia, Ajit Saravanan, Ishan Dave, Saadullah Ismail, Shikhar Shiromani, Emily Huang, Ruizhe Li, Kevin Zhu