arXiv Computer Vision

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

GazeFlow is a new framework for egocentric gaze prediction that models gaze as a joint distribution of temporal positions conditioned on both top‑down task cues and bottom‑up visual saliency. It employs conditional flow matching to iteratively transform Gaussian noise into realistic gaze trajectories, using a velocity field informed by video‑encoded visual features and global task queries. On standard benchmarks, GazeFlow outperforms existing methods on per‑frame metrics and produces trajectories that better reflect human gaze dynamics.

Hugging Face Trending Papers
Aug 11

Gaze Target Estimation Anywhere with Concepts

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis.

arXiv Computer Vision
Sep 17

GazeDiT: Gaze-Accurate Diffusion Image Generation for Eye Tracking via Spatial Conditioning

GazeDiT is a diffusion model that generates synthetic eye images conditioned on a precise 4‑dimensional gaze vector by using a spatial condition derived from pupil and iris geometry. During training, a frozen SegFormer extracts this geometry from real images, while inference employs a physical eye renderer to produce gaze‑consistent geometries without needing a source image. The model achieves lower tail gaze‑label error than other diffusion baselines and improves downstream eye‑tracking accuracy, reducing error on challenging cases from 3.05° to 2.80°.

By Dongze Wu, David Colmenares, Fengting Yang, Jogendra Nath Kundu, Yao Xie, Ali Behrooz, Conny Lu
arXiv Computer Vision
4d ago

OpenVAM: Open-World Visual Attention Modeling with VLMs

OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.

By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
Hugging Face Trending Papers
Aug 19

EgoHRV: Continuous Heart Rate Variability Estimation from Egocentric Systems for Autonomic Response and Skill Assessment

EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.