arXiv AI By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

Read the original on arXiv AI →

The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

Visual Attention Faithfulness in Vision-Language Models is Heterogeneous

The study investigates whether attention weights in Vision‑Language Models (VLMs) accurately reflect model reasoning for visual inputs. Using causal perturbation analysis, it identifies three distinct processing modes—Faithful‑Sufficient, Faithful‑Distributed, and Non‑Focal—indicating heterogeneous visual attention faithfulness. The research also shows that human‑annotated ground‑truth regions align with model attention in only about 60% of cases, highlighting a systematic divergence between model visual reliance and human intuition across VQA, document, and chart tasks.

By Xurui Song, Weishi Wang, Zhongqi Yue, Kuluhan Binici, Tao Bai, Hongxin Shao, Daniel Dahlmeier, Jun Luo
arXiv Machine Learning
6d ago

The Alignment Illusion in Multimodal Large Language Models

The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.

By Hong-Han Wang, Yuntao Wang, Hu Ding