arXiv Machine Learning By Israfel Salazar, Stella Frank, Dan Oneata, Desmond Elliott, Constanza Fierro

Pathways of Visual Information Flow in Vision-Language Models

Read the original on arXiv Machine Learning →

arXiv:2607. 03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 27

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

The paper demonstrates that vision‑language models (VLMs) possess a small set of attention heads, called Visual Retrieval Heads (VRHs), that are causally responsible for linking text prompts to specific image regions. By adapting head‑scoring techniques from language models, the authors identify VRHs as the heads whose attention from output prediction tokens, summed over the ground‑truth referent region, most reliably indicates causal grounding. Experiments across eleven VLMs and five referring‑expression benchmarks show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect, and that VRHs generalize across diverse visual tasks and transfer across models sharing an LLM backbone.

arXiv Computation and Language
Sep 7

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

The study examines how Vision‑Language Models (VLMs) integrate visual evidence into language‑based decisions by applying layer‑wise causal interventions on video‑text attention pathways in a video‑based generative multiple‑choice setting. Findings reveal that visual information is primarily incorporated while processing candidate answer options, with nouns serving as key semantic anchors and verbs becoming important during temporal reasoning. The research also uncovers a distinct pattern in temporal reasoning, indicating that VLMs struggle to reconstruct sequential information across video frames, possibly due to linguistic biases in temporal expressions.

By Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini, Albert Gatt
arXiv Computer Vision
Aug 28

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

The paper introduces Visual Retrieval Heads (VRHs), a small fraction of attention heads in vision‑language models that are causally responsible for grounding text descriptions to image regions. By recasting head‑scoring methods and evaluating across eleven VLMs and five benchmarks, the authors show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect. VRHs generalize across various visual reference tasks, preserve output format while corrupting localization, and transfer causally across models sharing an LLM backbone.

By Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung