arXiv:2605. 22775v2 Announce Type: replace-cross Abstract: Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such as driver vigilance monitoring or automated flight deck assistance, yet two challenges persist: handling frequent data missingness from blinks and tracking failures, and efficiently modeling long-range temporal dependencies.
By Amir Mousavi, Mohammad Sadegh Sirjani, Erfan Nourbakhsh, Mimi Xie, Rocky Slavin, Leslie Neely, John Davis, John Quarles
EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.
EyeMakeYou is a multi‑conditional denoising diffusion model that synthesizes high‑frequency, subject‑specific gaze velocity sequences. It conditions on identity, task, and self‑reported subjective states (difficulty, mental tiredness, eye tiredness) to generate realistic 5‑second, 1000‑Hz bivariate gaze data from a reference trajectory. Experiments on the GazeBase dataset show that EyeMakeYou outperforms existing generative methods in spatial accuracy and real‑synthetic similarity while preserving task‑dependent associations with subjective reports.
By Kamrul Hasan, Mehedi Hasan Raju, Oleg V. Komogortsev
AOI-Net introduces a structural face AOI-guided Eye‑Gaze Track Network that jointly models short‑term temporal dynamics and AOI‑level structural organization for Autism Spectrum Disorder detection. The network uses a gating mechanism to adaptively combine complementary representations and incorporates class‑distribution‑aware learning to address the imbalance between ASD and typically developing participants. Experiments on a large clinical eye‑tracking database with over 1,300 participants demonstrate that AOI‑Net outperforms state‑of‑the‑art methods and offers interpretable gaze‑behavior modeling for scalable AI‑driven ASD screening.
By Zhanpei Huang, Binbin Sun, Jialiang Chen, Yiou Wang, Taochen Chen, Yuzhu Ji, Yiqun Zhang, Yiu-Ming Cheung
UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.
By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez
arXiv:2607. 22721v1 Announce Type: cross Abstract: Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning.
By Nassira Ait Mehdi, Milissa Temmam, Slimane Larabi