The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li
arXiv:2609.05517v1 Announce Type: cross
Abstract: Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free vi...
By Han Zhang
arXiv:2606. 14703v1 Announce Type: cross Abstract: How a vision-language model internally solves the task of describing an image is far from obvious.
By Rohit Gandikota, David Bau
arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.
By Pengcheng Pan, Yonekura Shogo, Yasuo Kuniyosh
arXiv:2606. 17389v1 Announce Type: cross Abstract: Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical.
By Logan Mann, Yi Xia, Ajit Saravanan, Ishan Dave, Saadullah Ismail, Shikhar Shiromani, Emily Huang, Ruizhe Li, Kevin Zhu
arXiv:2607. 16165v1 Announce Type: cross Abstract: Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot.
By Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
arXiv:2609.07474v2 Announce Type: replace
Abstract: Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should...
By Peng Xie, Amr Alanwar
arXiv:2609.18011v1 Announce Type: new
Abstract: In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evide...
By Nan Li, Albert Gatt, Massimo Poesio
arXiv:2609.05522v1 Announce Type: cross
Abstract: Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of pr...
By Laxman Basnet, Alexander Szorkovszky, Pedro G. Lind, Anis Yazidi, Shailendra Bhandari
arXiv:2610.00922v1 Announce Type: new
Abstract: Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict...
By Jungmin Lee, Niamat Ullah, Yoseob Han
arXiv:2308. 06035v4 Announce Type: replace Abstract: Humans routinely draw on visual context to predict upcoming words.
By Viktor Kewenig, Andrew Lampinen, Samuel A. Nastase, Christopher Edwards, Quitterie Lacome D'Elascombe, Akilles Rechardt, Jeremy I Skipper, Gabriella Vigliocco
The paper introduces a perception interface that separates vision from language in vision‑language models. A frozen perception stack detects objects, a deterministic semantic serializer converts the perceived state into text, and a standard text‑only large language model (LLM) answers questions. Experiments on a campus‑robot benchmark show that this serialized interface outperforms a zero‑shot VLM of the same language‑model size, especially as the language model shrinks, and that the advantage persists under paraphrase and different supervision regimes.
By Cong Xu, Ravi Sankar