GazeFS is a model that predicts and stabilizes target‑centered gaze trajectories using a variable‑length gaze‑head history, without requiring target information during inference. It maps this history to the next target‑center direction and a short‑horizon Search/Focus estimate, improving focus target centering and reducing residual gaze error. Across 7,960 acquisition episodes from 30 participants, GazeFS reduces Focus episode bias, dispersion, and P90 target error by 0.182°, 0.257°, and 0.400°, respectively, while maintaining high phase‑balanced accuracy and AUPRC.
By Yaozheng Xia, Zaiping Zhu, Bo Pang, Minghao Xie, Hui Li, Shaorong Wang, Sheng Li
The paper presents a framework that uses saliency maps to create hierarchical attention profiles, tracking how deep reinforcement learning agents allocate attention over time. By comparing these attention trajectories across different conditions and linking them to behavioral metrics, the study reveals algorithm‑specific biases, unintended reward‑driven strategies, and overfitting to redundant sensory inputs. Experiments on Atari 2600 games, custom Pong environments, and biomechanical visuomotor simulations demonstrate that these attention patterns correspond to measurable behavioral differences, establishing attention trajectories as a diagnostic tool beyond traditional performance metrics.
By Charlotte Beylier, Hannah Selder, Arthur Fleig, Simon M. Hofmann, Nico Scherf
arXiv:2610.00922v1 Announce Type: new
Abstract: Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict...
By Jungmin Lee, Niamat Ullah, Yoseob Han
arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.
By Pengcheng Pan, Yonekura Shogo, Yasuo Kuniyosh
The paper introduces the Causal Context-Gated Forecaster (CCGF) for predicting a driver's gaze during dashboard-mounted tracker dropouts. CCGF uses a 60‑frame history of gaze and head pose combined with DINOv3 scene features, and a learned reliability gate adjusts the influence of these inputs as the dropout progresses. Experiments on 2,047 naturalistic driving events show that live scene updates reduce median error by 33% compared to history‑only forecasting, while frozen scene input yields higher error, demonstrating the value of real‑time scene information.
By Shabnam Shabani, Ghazal Farhani
GameWAM is the first World-Action Model designed for native closed-loop gameplay and GUI control in modern video games. It jointly generates future visual observations and executable keyboard-mouse trajectories using parallel visual and action generative processes, block-causal conditioning, and flow matching. The model predicts gameplay/GUI mode at each step, handles heterogeneous native controls, and employs block-cycle control for long-horizon interaction, achieving competitive task success with fewer native actions than prior agents.
By Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li
arXiv:2606. 12200v1 Announce Type: cross Abstract: We study policy representation learning from unlabeled multi-policy behavioral data.
By Andrew Kang, Priya Narasimhan
The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts.
"whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."
By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
arXiv:2607. 04334v1 Announce Type: new Abstract: Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such as a DOM or accessibility tree.
By Guijia Zhang, Harry Yang
arXiv:2607. 21290v1 Announce Type: cross Abstract: Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as modern game telemetry provides multiple related supervision signals from the same structured observations.
By Jonas Pech\'e, Aliaksei Tsishurou, Alexander Zap, G\"unter Wallner
arXiv:2610.03276v1 Announce Type: new
Abstract: Video saliency prediction is inherently harder to model than static image saliency due to the additional temporal dimension. Video saliency benchmarks...
By Susmit Agrawal, Rebecca Wanner, Juliane Verwiebe, Matthias Tangemann, Matthias Bethge, Matthias K\"ummerer