arXiv AI

Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision

arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.

arXiv AI
4d ago

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.

By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li
arXiv Computer Vision
6d ago

OpenVAM: Open-World Visual Attention Modeling with VLMs

OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.

By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
arXiv Computer Vision
3d ago

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

GazeFlow is a new framework for egocentric gaze prediction that models gaze as a joint distribution of temporal positions conditioned on both top‑down task cues and bottom‑up visual saliency. It employs conditional flow matching to iteratively transform Gaussian noise into realistic gaze trajectories, using a velocity field informed by video‑encoded visual features and global task queries. On standard benchmarks, GazeFlow outperforms existing methods on per‑frame metrics and produces trajectories that better reflect human gaze dynamics.

By Sheng Zhao, Weikai Lin, Yuhao Zhu
Hugging Face Trending Papers
Aug 11

Gaze Target Estimation Anywhere with Concepts

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis.