arXiv:2609.05522v1 Announce Type: cross
Abstract: Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of pr...
By Laxman Basnet, Alexander Szorkovszky, Pedro G. Lind, Anis Yazidi, Shailendra Bhandari
arXiv:2608.29739v1 Announce Type: new
Abstract: Geometric eye trackers can provide the spatial accuracy required for gaze-based interaction and multimodal studies, but their measurements remain sensi...
By Jiaqi Liu, Zixuan Wang, Yuhong Zhang, Dingkang Liang, Jane Hanqi Li, Tzyy-Ping Jung, Gert Cauwenberghs
EyeMakeYou is a multi‑conditional denoising diffusion model that synthesizes high‑frequency, subject‑specific gaze velocity sequences. It conditions on identity, task, and self‑reported subjective states (difficulty, mental tiredness, eye tiredness) to generate realistic 5‑second, 1000‑Hz bivariate gaze data from a reference trajectory. Experiments on the GazeBase dataset show that EyeMakeYou outperforms existing generative methods in spatial accuracy and real‑synthetic similarity while preserving task‑dependent associations with subjective reports.
By Kamrul Hasan, Mehedi Hasan Raju, Oleg V. Komogortsev
arXiv:2507. 15833v3 Announce Type: replace-cross Abstract: Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing.
By Ian Chuang, Jinyu Zou, Andrew Lee, Dechen Gao, Iman Soltani
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis.
GazeFS is a model that predicts and stabilizes target‑centered gaze trajectories using a variable‑length gaze‑head history, without requiring target information during inference. It maps this history to the next target‑center direction and a short‑horizon Search/Focus estimate, improving focus target centering and reducing residual gaze error. Across 7,960 acquisition episodes from 30 participants, GazeFS reduces Focus episode bias, dispersion, and P90 target error by 0.182°, 0.257°, and 0.400°, respectively, while maintaining high phase‑balanced accuracy and AUPRC.
By Yaozheng Xia, Zaiping Zhu, Bo Pang, Minghao Xie, Hui Li, Shaorong Wang, Sheng Li
arXiv:2608. 11367v1 Announce Type: cross Abstract: Estimating human gaze targets from images in-the-wild is an important and formidable task.
By Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg
arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.
By Pengcheng Pan, Yonekura Shogo, Yasuo Kuniyosh
arXiv:2606. 14703v1 Announce Type: cross Abstract: How a vision-language model internally solves the task of describing an image is far from obvious.
By Rohit Gandikota, David Bau
The paper introduces EyeControl, an intent-driven image retouching agent that enhances visual focus by guiding attention to a target region with minimal user input. It combines a multi‑modal large language model to interpret user intent and a diffusion‑based retouching executor that aligns its attention map with a pseudo‑intent map, while an operation‑consistency constraint ensures natural global and local adjustments. The authors also present ControlArt‑Bench, a dataset for evaluating visual focus enhancement, and demonstrate that EyeControl achieves perceptually appealing results with stronger intent alignment.
By Chujie Qin, Zilong Zhang, Zewei Chang, Chunle Guo, Ruixing Wang, Tao Hu, Ming-Ming Cheng, Chongyi Li
arXiv:2608. 15614v1 Announce Type: cross Abstract: The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget.
By Matteo Stoiber, Niels Buus Lassen
arXiv:2603. 06697v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks.
By Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu, Xiang Li, Quanzheng Li, Tianming Liu, Lin Zhao