arXiv AI By Yi Xia, Ibrahim Khan, Mury Fajar Dewantoro, Wenwen Ouyang, Ruck Thawonmas

Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
5d ago

Do Vision Language Models Understand Human Engagement in Games?

The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts. "whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."

By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
arXiv Computer Vision
Sep 3

Video2Reaction: Training Foundation Video Models to Predict Audience Reaction

Video2Reaction is a multimodal dataset that links short movie segments to the emotional reactions of viewers, gathered from social media comments. The dataset models reactions as distributions over categorical emotions, capturing the subjective and ambiguous nature of emotional perception. Experiments show that vision‑language models fine‑tuned with LoRA learn effectively from Video2Reaction and outperform specialized baselines, and that models pre‑fine‑tuned on this dataset transfer well to other emotion prediction tasks.

By Sidong Zhang, Trang Nguyen, Shiv Shankar, Gauri Jagatap, Deepak Chandran, Andrea Fanelli, Madalina Fiterau
arXiv Computer Vision
Sep 21

Learned Parametric Emotion Editing: Real-Time Affective Filtering for On-Device Social Media Video

The paper presents a real‑time on‑device system for editing the emotional intensity of visual content. Using a MobileNetV4 backbone with FiLM‑based conditioning, the model predicts parameters for differentiable global transformations in a single 3.7 ms forward pass, replacing 80‑second per‑image optimization. A user study with 54 participants showed reduced viewer arousal and higher perceived quality compared to a grayscale filter, and the system runs at 60 fps on a Samsung Galaxy S23.

By Musa Rochi, Marcel Schubert, Christoph Gebhardt
arXiv Computer Vision
2d ago

Moving Forward with Video Saliency: A New Dataset and Benchmark where Motion Matters

The paper introduces SalTempto, a new video saliency benchmark featuring 224 highly dynamic clips from HACS‑Segments, each with event context and gaze data from up to 16 subjects. It shows that static baselines perform poorly on SalTempto compared to LEDOV, while fine‑tuned temporal models achieve significant gains, yet still leave a large portion of explainable gaze information unaccounted for. The dataset highlights the need for better temporal modeling in video saliency research.

By Susmit Agrawal, Rebecca Wanner, Juliane Verwiebe, Matthias Tangemann, Matthias Bethge, Matthias K\"ummerer