arXiv:2608.22926v1 Announce Type: new
Abstract: Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw trac...
By Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci
EyeMakeYou is a multi‑conditional denoising diffusion model that synthesizes high‑frequency, subject‑specific gaze velocity sequences. It conditions on identity, task, and self‑reported subjective states (difficulty, mental tiredness, eye tiredness) to generate realistic 5‑second, 1000‑Hz bivariate gaze data from a reference trajectory. Experiments on the GazeBase dataset show that EyeMakeYou outperforms existing generative methods in spatial accuracy and real‑synthetic similarity while preserving task‑dependent associations with subjective reports.
By Kamrul Hasan, Mehedi Hasan Raju, Oleg V. Komogortsev
arXiv:2609.17814v1 Announce Type: new
Abstract: Diffusion models are increasingly used to generate synthetic training data, but precise label control remains difficult when the conditioning signal is...
By Dongze Wu, David Colmenares, Fengting Yang, Jogendra Nath Kundu, Yao Xie, Ali Behrooz, Conny Lu
arXiv:2608. 16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath.
By Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno
GazeFS is a model that predicts and stabilizes target‑centered gaze trajectories using a variable‑length gaze‑head history, without requiring target information during inference. It maps this history to the next target‑center direction and a short‑horizon Search/Focus estimate, improving focus target centering and reducing residual gaze error. Across 7,960 acquisition episodes from 30 participants, GazeFS reduces Focus episode bias, dispersion, and P90 target error by 0.182°, 0.257°, and 0.400°, respectively, while maintaining high phase‑balanced accuracy and AUPRC.
By Yaozheng Xia, Zaiping Zhu, Bo Pang, Minghao Xie, Hui Li, Shaorong Wang, Sheng Li
arXiv:2606. 14703v1 Announce Type: cross Abstract: How a vision-language model internally solves the task of describing an image is far from obvious.
By Rohit Gandikota, David Bau
The paper introduces the Causal Context-Gated Forecaster (CCGF) for predicting a driver's gaze during dashboard-mounted tracker dropouts. CCGF uses a 60‑frame history of gaze and head pose combined with DINOv3 scene features, and a learned reliability gate adjusts the influence of these inputs as the dropout progresses. Experiments on 2,047 naturalistic driving events show that live scene updates reduce median error by 33% compared to history‑only forecasting, while frozen scene input yields higher error, demonstrating the value of real‑time scene information.
By Shabnam Shabani, Ghazal Farhani
arXiv:2608.29739v1 Announce Type: new
Abstract: Geometric eye trackers can provide the spatial accuracy required for gaze-based interaction and multimodal studies, but their measurements remain sensi...
By Jiaqi Liu, Zixuan Wang, Yuhong Zhang, Dingkang Liang, Jane Hanqi Li, Tzyy-Ping Jung, Gert Cauwenberghs
arXiv:2608. 15614v1 Announce Type: cross Abstract: The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget.
By Matteo Stoiber, Niels Buus Lassen
arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.
By Pengcheng Pan, Yonekura Shogo, Yasuo Kuniyosh
arXiv:2606. 12987v1 Announce Type: cross Abstract: Action-conditioned world models let an autonomous vehicle predict future camera scenes from its own planned controls, enabling planning and simulation without real-world rollouts, but at compact, trainable scale the futures are ambiguous and the field's standard distortion metrics actively mislead: they reward a blurry regression mean over a realistic prediction.
By Ruslan Sharifullin, Benjamin Jiang, Kai Xi Chew
GlanceWAM introduces a sparse test‑time imagination approach for world‑action models that decouples visual imagination from control. By asynchronously generating a single lookahead frame on a slow clock and decoding action chunks at a 48 ms control rate purely in latent space, it avoids latency while maintaining high success. The method achieves 72.2 % on the RoboCasa kitchen benchmark and 99.0 % on LIBERO, running 24× faster than synchronous baselines.
By Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Zichong Yang, Yinlin Chen, Lifeng Zhou, Chang-Tien Lu