arXiv AI

Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage

arXiv AI
5d ago

Do Vision Language Models Understand Human Engagement in Games?

The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts. "whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."

By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
arXiv Computer Vision
Sep 3

Video2Reaction: Training Foundation Video Models to Predict Audience Reaction

Video2Reaction is a multimodal dataset that links short movie segments to the emotional reactions of viewers, gathered from social media comments. The dataset models reactions as distributions over categorical emotions, capturing the subjective and ambiguous nature of emotional perception. Experiments show that vision‑language models fine‑tuned with LoRA learn effectively from Video2Reaction and outperform specialized baselines, and that models pre‑fine‑tuned on this dataset transfer well to other emotion prediction tasks.

By Sidong Zhang, Trang Nguyen, Shiv Shankar, Gauri Jagatap, Deepak Chandran, Andrea Fanelli, Madalina Fiterau
arXiv Computer Vision
Sep 21

Learned Parametric Emotion Editing: Real-Time Affective Filtering for On-Device Social Media Video

The paper presents a real‑time on‑device system for editing the emotional intensity of visual content. Using a MobileNetV4 backbone with FiLM‑based conditioning, the model predicts parameters for differentiable global transformations in a single 3.7 ms forward pass, replacing 80‑second per‑image optimization. A user study with 54 participants showed reduced viewer arousal and higher perceived quality compared to a grayscale filter, and the system runs at 60 fps on a Samsung Galaxy S23.

By Musa Rochi, Marcel Schubert, Christoph Gebhardt
arXiv Computer Vision
2d ago

Moving Forward with Video Saliency: A New Dataset and Benchmark where Motion Matters

The paper introduces SalTempto, a new video saliency benchmark featuring 224 highly dynamic clips from HACS‑Segments, each with event context and gaze data from up to 16 subjects. It shows that static baselines perform poorly on SalTempto compared to LEDOV, while fine‑tuned temporal models achieve significant gains, yet still leave a large portion of explainable gaze information unaccounted for. The dataset highlights the need for better temporal modeling in video saliency research.

By Susmit Agrawal, Rebecca Wanner, Juliane Verwiebe, Matthias Tangemann, Matthias Bethge, Matthias K\"ummerer
arXiv Computer Vision
Sep 21

OpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection

OpenSAL360 is an open‑source platform that enables scalable, low‑cost collection of 360° video saliency data using only a standard screen, mouse, and internet connection. It bypasses the need for VR headsets, allowing parallel data collection from crowdsourced assessors. The authors validated the protocol against seven VR eye‑tracking datasets, performed ablation studies, and released a new dataset of 500 omnidirectional videos annotated by over 2,000 assessors, the largest in the field to date.

By Alexey Bryncev, Andrey Moskalenko, Kira Shilovskaya, Ivan Kosmynin, Dmitriy Vatolin
arXiv Computer Vision
Sep 16

RefGlitch-Bench: A Benchmark for Reference-based Gameplay Glitch Detection with Vision-Language Models

arXiv:2604.11082v2 Announce Type: replace Abstract: Visual glitches in video games degrade player experience and perceived quality, yet manual quality assurance cannot keep pace with the growing test...

By Yakun Yu, Ashley Wiens, Adri\'an Barahona-R\'ios, Benedict Wilkins, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer