arXiv Machine Learning

Consensus Clustering of Free-Viewing Gaze Data: New Insights into Human-Information Interaction

arXiv:2606. 30035v1 Announce Type: cross Abstract: Free-viewing gaze data provides a rich, task-free window into human visual attention.

arXiv Computer Vision
Sep 7

EyeMakeYou: Identity-, Task-, and Subjective-State-Conditioned Diffusion for High-Frequency Gaze Synthesis

EyeMakeYou is a multi‑conditional denoising diffusion model that synthesizes high‑frequency, subject‑specific gaze velocity sequences. It conditions on identity, task, and self‑reported subjective states (difficulty, mental tiredness, eye tiredness) to generate realistic 5‑second, 1000‑Hz bivariate gaze data from a reference trajectory. Experiments on the GazeBase dataset show that EyeMakeYou outperforms existing generative methods in spatial accuracy and real‑synthetic similarity while preserving task‑dependent associations with subjective reports.

By Kamrul Hasan, Mehedi Hasan Raju, Oleg V. Komogortsev
arXiv AI
Sep 1

AOI-Net: Structural Face AOI-Guided Eye-Gaze Track Representation Learning for Autism Spectrum Disorder Detection

AOI-Net introduces a structural face AOI-guided Eye‑Gaze Track Network that jointly models short‑term temporal dynamics and AOI‑level structural organization for Autism Spectrum Disorder detection. The network uses a gating mechanism to adaptively combine complementary representations and incorporates class‑distribution‑aware learning to address the imbalance between ASD and typically developing participants. Experiments on a large clinical eye‑tracking database with over 1,300 participants demonstrate that AOI‑Net outperforms state‑of‑the‑art methods and offers interpretable gaze‑behavior modeling for scalable AI‑driven ASD screening.

By Zhanpei Huang, Binbin Sun, Jialiang Chen, Yiou Wang, Taochen Chen, Yuzhu Ji, Yiqun Zhang, Yiu-Ming Cheung
Hugging Face Trending Papers
Aug 11

Gaze Target Estimation Anywhere with Concepts

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis.

arXiv Computer Vision
6d ago

OpenVAM: Open-World Visual Attention Modeling with VLMs

OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.

By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
arXiv Machine Learning
4d ago

Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence

The paper introduces the concept of protocol divergence, showing that identical nominal missing rates can lead to vastly different learning regimes in incomplete multi‑view clustering. It critiques existing evaluation practices that ignore observation structure and proposes CRAFT, a train‑once framework that fuses observed views with mask‑aware attention, enabling efficient deployment across multiple missing‑view protocols. Experiments on CUB, MultiFashion, and other benchmarks demonstrate CRAFT’s superior performance and significant computational savings through checkpoint reuse.

By Haolu Liu, Xiyue Wang, Xuanting Xie, Liangjian Wen, Zhao Kang