The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts.
"whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."
By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
Video2Reaction is a multimodal dataset that links short movie segments to the emotional reactions of viewers, gathered from social media comments. The dataset models reactions as distributions over categorical emotions, capturing the subjective and ambiguous nature of emotional perception. Experiments show that vision‑language models fine‑tuned with LoRA learn effectively from Video2Reaction and outperform specialized baselines, and that models pre‑fine‑tuned on this dataset transfer well to other emotion prediction tasks.
By Sidong Zhang, Trang Nguyen, Shiv Shankar, Gauri Jagatap, Deepak Chandran, Andrea Fanelli, Madalina Fiterau
arXiv:2608.24680v1 Announce Type: new
Abstract: Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in-the...
By Wenxuan Shen, Dongna Jin, Dongping Chen
arXiv:2609.36563v1 Announce Type: new
Abstract: Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible...
By Yihao Qian, Runhao Zeng, Sicheng Zhao, Feng Liang, Hongmin Cai, Mingkui Tan
The paper presents a real‑time on‑device system for editing the emotional intensity of visual content. Using a MobileNetV4 backbone with FiLM‑based conditioning, the model predicts parameters for differentiable global transformations in a single 3.7 ms forward pass, replacing 80‑second per‑image optimization. A user study with 54 participants showed reduced viewer arousal and higher perceived quality compared to a grayscale filter, and the system runs at 60 fps on a Samsung Galaxy S23.
By Musa Rochi, Marcel Schubert, Christoph Gebhardt
The paper introduces SalTempto, a new video saliency benchmark featuring 224 highly dynamic clips from HACS‑Segments, each with event context and gaze data from up to 16 subjects. It shows that static baselines perform poorly on SalTempto compared to LEDOV, while fine‑tuned temporal models achieve significant gains, yet still leave a large portion of explainable gaze information unaccounted for. The dataset highlights the need for better temporal modeling in video saliency research.
By Susmit Agrawal, Rebecca Wanner, Juliane Verwiebe, Matthias Tangemann, Matthias Bethge, Matthias K\"ummerer
OpenSAL360 is an open‑source platform that enables scalable, low‑cost collection of 360° video saliency data using only a standard screen, mouse, and internet connection. It bypasses the need for VR headsets, allowing parallel data collection from crowdsourced assessors. The authors validated the protocol against seven VR eye‑tracking datasets, performed ablation studies, and released a new dataset of 500 omnidirectional videos annotated by over 2,000 assessors, the largest in the field to date.
By Alexey Bryncev, Andrey Moskalenko, Kira Shilovskaya, Ivan Kosmynin, Dmitriy Vatolin
arXiv:2508.08966v2 Announce Type: replace
Abstract: The attention mechanism lies at the core of the transformer architecture, providing an interpretable model-internal signal that has motivated a gro...
By Marte Eggen, Jacob Lysn{\ae}s-Larsen, Inga Str\"umke
arXiv:2604.11082v2 Announce Type: replace
Abstract: Visual glitches in video games degrade player experience and perceived quality, yet manual quality assurance cannot keep pace with the growing test...
By Yakun Yu, Ashley Wiens, Adri\'an Barahona-R\'ios, Benedict Wilkins, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer
arXiv:2607. 06875v1 Announce Type: cross Abstract: Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis.
By Trang Nguyen, Sidong Zhang, Shiv Shankar, Gauri Jagatap, Deepak Chandran, Andrea Fanelli, Madalina Fiterau
arXiv:2607. 10165v1 Announce Type: cross Abstract: Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion.
By Dexiang Hong, Yijie Guo, Weidong Chen, Xinyan Liu, Zixuan Zou, Zhendong Mao, Yongdong Zhang
arXiv:2609.22947v1 Announce Type: new
Abstract: Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, exist...
By Zhenchen Tang, Yang Li, Songlin Yang, Bo Peng, Xiaotong Zhao, Shuai Li, Haotian Fan, Alan Zhao, Jing Dong