arXiv:2608.22926v1 Announce Type: new
Abstract: Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw trac...
By Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci
arXiv:2609.05522v1 Announce Type: cross
Abstract: Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of pr...
By Laxman Basnet, Alexander Szorkovszky, Pedro G. Lind, Anis Yazidi, Shailendra Bhandari
arXiv:2606. 25177v1 Announce Type: new Abstract: Cognitive workload monitoring is important for adaptive rehabilitation and assistive interfaces, where task difficulty, pacing, and feedback should be adjusted according to the user's cognitive state to avoid overload and under-challenge.
By Guorui Lu, Shaohua Guan, Zhen Xu, Qinyu Chen
arXiv:2608. 03025v1 Announce Type: new Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence.
By Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
arXiv:2608. 03025v3 Announce Type: replace Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence.
By Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
AOI-Net introduces a structural face AOI-guided Eye‑Gaze Track Network that jointly models short‑term temporal dynamics and AOI‑level structural organization for Autism Spectrum Disorder detection. The network uses a gating mechanism to adaptively combine complementary representations and incorporates class‑distribution‑aware learning to address the imbalance between ASD and typically developing participants. Experiments on a large clinical eye‑tracking database with over 1,300 participants demonstrate that AOI‑Net outperforms state‑of‑the‑art methods and offers interpretable gaze‑behavior modeling for scalable AI‑driven ASD screening.
By Zhanpei Huang, Binbin Sun, Jialiang Chen, Yiou Wang, Taochen Chen, Yuzhu Ji, Yiqun Zhang, Yiu-Ming Cheung
arXiv:2606. 30035v1 Announce Type: cross Abstract: Free-viewing gaze data provides a rich, task-free window into human visual attention.
By Beryl Gnanaraj, Jaya Sreevalsan-Nair, Saqib Alam Ansari, Maanasa Rajaraman
arXiv:2608. 11367v1 Announce Type: cross Abstract: Estimating human gaze targets from images in-the-wild is an important and formidable task.
By Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg
The paper tackles the challenge of predicting student engagement from online tutoring videos, noting that engagement is a complex, multidimensional construct influenced by behavioral, emotional, and cognitive states. By analyzing the CASED dataset, the authors highlight the difficulty posed by high inter‑person variability and subjective annotations. They propose a multimodal framework that fuses implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral cues such as head pose, gaze, facial action units, emotion, and wavelet‑based audio features, integrating them via a Perceiver IO bottleneck and modeling participant personalities with variational posteriors. The system employs evidential regression and spectral‑normalized Gaussian process classification heads to provide uncertainty‑aware predictions, achieving competitive performance on the CASED challenge test set while offering well‑calibrated uncertainty metrics.
By Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.
arXiv:2607. 29337v1 Announce Type: cross Abstract: Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of annotated data in medical imaging, particularly in optical coherence tomography (OCT) of mouse eyes, where manual retinal layer delineation is labour-intensive due to tiny structures and required expertise, resulting in scarce datasets.
By Fernando Garc\'ia-Torres, Roc\'io del Amor, Sandra Morales, \'Alvaro Barroso, Peter Heiduschka, Bj\"orn Kemper, Valery Naranjo
arXiv:2609.17814v1 Announce Type: new
Abstract: Diffusion models are increasingly used to generate synthetic training data, but precise label control remains difficult when the conditioning signal is...
By Dongze Wu, David Colmenares, Fengting Yang, Jogendra Nath Kundu, Yao Xie, Ali Behrooz, Conny Lu