Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
arXiv:2608. 16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath.
The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
arXiv:2608. 16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath.
arXiv:2609.05517v1 Announce Type: cross Abstract: Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free vi...
arXiv:2308. 06035v4 Announce Type: replace Abstract: Humans routinely draw on visual context to predict upcoming words.
arXiv:2606. 14703v1 Announce Type: cross Abstract: How a vision-language model internally solves the task of describing an image is far from obvious.
The study investigates whether attention weights in Vision‑Language Models (VLMs) accurately reflect model reasoning for visual inputs. Using causal perturbation analysis, it identifies three distinct processing modes—Faithful‑Sufficient, Faithful‑Distributed, and Non‑Focal—indicating heterogeneous visual attention faithfulness. The research also shows that human‑annotated ground‑truth regions align with model attention in only about 60% of cases, highlighting a systematic divergence between model visual reliance and human intuition across VQA, document, and chart tasks.
The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.
arXiv:2606. 31054v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image.
arXiv:2609.18011v1 Announce Type: new Abstract: In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evide...
The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects cross‑modal content integration. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that common scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from intact visual tokens, a phenomenon they term the "alignment illusion." They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task accuracy and reveals when internal geometry diverges from performance.
The study investigates whether multimodal large language models (MLLMs) report bistable images, like the duck‑rabbit, in a manner similar to humans. Using the LLaVA family, researchers examined two dimensions: modulability (the influence of visual cues and linguistic priors) and exclusivity (whether responses commit to a single interpretation). Results show that both visual and linguistic manipulations shift reports in human‑consistent ways while maintaining predominantly exclusive responses, driven by competing image‑token representations and distinct bottom‑up and top‑down pathways.
arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.
The paper shows that the wording of prompts in vision‑language models (VLMs) can either improve or worsen robustness to image corruption. Verbose prompts broaden the cross‑modal attention’s frequency filter, making the model less sensitive to corruptions, while semantically complex prompts narrow the filter and increase vulnerability. Experiments on Qwen3‑VL and LLaVA‑OneVision confirm that adding padding or verbose phrasing reduces answer drift by 70–81% on 8B models.