The paper introduces the Causal Context-Gated Forecaster (CCGF) for predicting a driver's gaze during dashboard-mounted tracker dropouts. CCGF uses a 60‑frame history of gaze and head pose combined with DINOv3 scene features, and a learned reliability gate adjusts the influence of these inputs as the dropout progresses. Experiments on 2,047 naturalistic driving events show that live scene updates reduce median error by 33% compared to history‑only forecasting, while frozen scene input yields higher error, demonstrating the value of real‑time scene information.
By Shabnam Shabani, Ghazal Farhani
GazeFS is a model that predicts and stabilizes target‑centered gaze trajectories using a variable‑length gaze‑head history, without requiring target information during inference. It maps this history to the next target‑center direction and a short‑horizon Search/Focus estimate, improving focus target centering and reducing residual gaze error. Across 7,960 acquisition episodes from 30 participants, GazeFS reduces Focus episode bias, dispersion, and P90 target error by 0.182°, 0.257°, and 0.400°, respectively, while maintaining high phase‑balanced accuracy and AUPRC.
By Yaozheng Xia, Zaiping Zhu, Bo Pang, Minghao Xie, Hui Li, Shaorong Wang, Sheng Li
arXiv:2608.22926v1 Announce Type: new
Abstract: Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw trac...
By Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci
GazeFlow is a new framework for egocentric gaze prediction that models gaze as a joint distribution of temporal positions conditioned on both top‑down task cues and bottom‑up visual saliency. It employs conditional flow matching to iteratively transform Gaussian noise into realistic gaze trajectories, using a velocity field informed by video‑encoded visual features and global task queries. On standard benchmarks, GazeFlow outperforms existing methods on per‑frame metrics and produces trajectories that better reflect human gaze dynamics.
By Sheng Zhao, Weikai Lin, Yuhao Zhu
arXiv:2609.05522v1 Announce Type: cross
Abstract: Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of pr...
By Laxman Basnet, Alexander Szorkovszky, Pedro G. Lind, Anis Yazidi, Shailendra Bhandari
arXiv:2608.29739v1 Announce Type: new
Abstract: Geometric eye trackers can provide the spatial accuracy required for gaze-based interaction and multimodal studies, but their measurements remain sensi...
By Jiaqi Liu, Zixuan Wang, Yuhong Zhang, Dingkang Liang, Jane Hanqi Li, Tzyy-Ping Jung, Gert Cauwenberghs
arXiv:2608. 15614v1 Announce Type: cross Abstract: The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget.
By Matteo Stoiber, Niels Buus Lassen
arXiv:2606. 14703v1 Announce Type: cross Abstract: How a vision-language model internally solves the task of describing an image is far from obvious.
By Rohit Gandikota, David Bau
EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.
arXiv:2608. 08947v1 Announce Type: cross Abstract: Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition.
By Lennox Anderson, Ahmed Boutar, Jonah Mulcrone, Tal Erez
GazeDiT is a diffusion model that generates synthetic eye images conditioned on a precise 4‑dimensional gaze vector by using a spatial condition derived from pupil and iris geometry. During training, a frozen SegFormer extracts this geometry from real images, while inference employs a physical eye renderer to produce gaze‑consistent geometries without needing a source image. The model achieves lower tail gaze‑label error than other diffusion baselines and improves downstream eye‑tracking accuracy, reducing error on challenging cases from 3.05° to 2.80°.
By Dongze Wu, David Colmenares, Fengting Yang, Jogendra Nath Kundu, Yao Xie, Ali Behrooz, Conny Lu
arXiv:2606. 25177v1 Announce Type: new Abstract: Cognitive workload monitoring is important for adaptive rehabilitation and assistive interfaces, where task difficulty, pacing, and feedback should be adjusted according to the user's cognitive state to avoid overload and under-challenge.
By Guorui Lu, Shaohua Guan, Zhen Xu, Qinyu Chen