Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts. "whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."
Video2Reaction is a multimodal dataset that links short movie segments to the emotional reactions of viewers, gathered from social media comments. The dataset models reactions as distributions over categorical emotions, capturing the subjective and ambiguous nature of emotional perception. Experiments show that vision‑language models fine‑tuned with LoRA learn effectively from Video2Reaction and outperform specialized baselines, and that models pre‑fine‑tuned on this dataset transfer well to other emotion prediction tasks.
arXiv:2608.24680v1 Announce Type: new Abstract: Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in-the...
arXiv:2609.36563v1 Announce Type: new Abstract: Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible...
The paper presents a real‑time on‑device system for editing the emotional intensity of visual content. Using a MobileNetV4 backbone with FiLM‑based conditioning, the model predicts parameters for differentiable global transformations in a single 3.7 ms forward pass, replacing 80‑second per‑image optimization. A user study with 54 participants showed reduced viewer arousal and higher perceived quality compared to a grayscale filter, and the system runs at 60 fps on a Samsung Galaxy S23.
The paper introduces SalTempto, a new video saliency benchmark featuring 224 highly dynamic clips from HACS‑Segments, each with event context and gaze data from up to 16 subjects. It shows that static baselines perform poorly on SalTempto compared to LEDOV, while fine‑tuned temporal models achieve significant gains, yet still leave a large portion of explainable gaze information unaccounted for. The dataset highlights the need for better temporal modeling in video saliency research.