Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
arXiv:2608. 16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath.
arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.
arXiv:2608. 16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath.
arXiv:2507. 15833v3 Announce Type: replace-cross Abstract: Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing.
arXiv:2609.05517v1 Announce Type: cross Abstract: Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free vi...
The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.
arXiv:2610.00922v1 Announce Type: new Abstract: Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict...
arXiv:2608.29739v1 Announce Type: new Abstract: Geometric eye trackers can provide the spatial accuracy required for gaze-based interaction and multimodal studies, but their measurements remain sensi...
GazeFlow is a new framework for egocentric gaze prediction that models gaze as a joint distribution of temporal positions conditioned on both top‑down task cues and bottom‑up visual saliency. It employs conditional flow matching to iteratively transform Gaussian noise into realistic gaze trajectories, using a velocity field informed by video‑encoded visual features and global task queries. On standard benchmarks, GazeFlow outperforms existing methods on per‑frame metrics and produces trajectories that better reflect human gaze dynamics.
arXiv:2608.22926v1 Announce Type: new Abstract: Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw trac...
arXiv:2606. 30035v1 Announce Type: cross Abstract: Free-viewing gaze data provides a rich, task-free window into human visual attention.
arXiv:2608. 11367v1 Announce Type: cross Abstract: Estimating human gaze targets from images in-the-wild is an important and formidable task.
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis.