arXiv Machine Learning

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games

arXiv:2604. 04439v2 Announce Type: replace Abstract: We study how different visual information sources contribute to human decision making in dynamic visual environments.

arXiv AI
Sep 4

GazeFS: Target-Centered Gaze-Trajectory Forecasting and Stabilization from Gaze-Head History

GazeFS is a model that predicts and stabilizes target‑centered gaze trajectories using a variable‑length gaze‑head history, without requiring target information during inference. It maps this history to the next target‑center direction and a short‑horizon Search/Focus estimate, improving focus target centering and reducing residual gaze error. Across 7,960 acquisition episodes from 30 participants, GazeFS reduces Focus episode bias, dispersion, and P90 target error by 0.182°, 0.257°, and 0.400°, respectively, while maintaining high phase‑balanced accuracy and AUPRC.

By Yaozheng Xia, Zaiping Zhu, Bo Pang, Minghao Xie, Hui Li, Shaorong Wang, Sheng Li
arXiv Machine Learning
Sep 4

Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning

The paper presents a framework that uses saliency maps to create hierarchical attention profiles, tracking how deep reinforcement learning agents allocate attention over time. By comparing these attention trajectories across different conditions and linking them to behavioral metrics, the study reveals algorithm‑specific biases, unintended reward‑driven strategies, and overfitting to redundant sensory inputs. Experiments on Atari 2600 games, custom Pong environments, and biomechanical visuomotor simulations demonstrate that these attention patterns correspond to measurable behavioral differences, establishing attention trajectories as a diagnostic tool beyond traditional performance metrics.

By Charlotte Beylier, Hannah Selder, Arthur Fleig, Simon M. Hofmann, Nico Scherf
arXiv Computer Vision
Sep 14

Context-Aware Causal Gaze Forecasting for Human-Vehicle Interaction During In-Cabin Tracking Dropouts

The paper introduces the Causal Context-Gated Forecaster (CCGF) for predicting a driver's gaze during dashboard-mounted tracker dropouts. CCGF uses a 60‑frame history of gaze and head pose combined with DINOv3 scene features, and a learned reliability gate adjusts the influence of these inputs as the dropout progresses. Experiments on 2,047 naturalistic driving events show that live scene updates reduce median error by 33% compared to history‑only forecasting, while frozen scene input yields higher error, demonstrating the value of real‑time scene information.

By Shabnam Shabani, Ghazal Farhani
arXiv AI
Aug 28

GameWAM: A World Action Model for Video Games

GameWAM is the first World-Action Model designed for native closed-loop gameplay and GUI control in modern video games. It jointly generates future visual observations and executable keyboard-mouse trajectories using parallel visual and action generative processes, block-causal conditioning, and flow matching. The model predicts gameplay/GUI mode at each step, handles heterogeneous native controls, and employs block-cycle control for long-horizon interaction, achieving competitive task success with fewer native actions than prior agents.

By Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
arXiv AI
6d ago

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.

By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li
arXiv AI
4d ago

Do Vision Language Models Understand Human Engagement in Games?

The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts. "whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."

By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu