To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD. The proposed framework leverages physiological signals as privileged information during training to guide a video-based student model in learning deep affective representations, while relying solely on non-intrusive video inputs at inference time.
The paper introduces TTSD‑FAR, a test‑time self‑distillation framework that adapts large video‑language models to missing‑modality scenarios in emotion recognition. A frozen teacher trained on complete modalities guides a low‑rank student, while Fisher‑Anchored Restoration monitors Fisher information to prevent drift and restore the student when distribution shifts occur. Experiments on MELD, DFEW, and BAH with up to 50% missing modalities show TTSD‑FAR consistently outperforms entropy‑based adaptation, retrieval‑augmented generation, and perplexity‑based generation, maintaining performance over long adaptation horizons.
By Muhammad Haseeb Aslam, Alessandro Koerich, Marco Pedersoli, Ali Etemad, Eric Granger
arXiv:2604. 15336v2 Announce Type: replace-cross Abstract: Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone.
By Shuangquan Feng, Laura Fleig, Ruisen Tu, Philip Chi, Edmund Bu, Melinda Ozel, Junhua Ma, Teng Fei, Virginia R. de Sa
arXiv:2608. 14675v1 Announce Type: cross Abstract: While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as photoplethysmography (PPG), its suitability for highly subjective tasks remains unproven.
By Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski
The paper tackles the challenge of predicting student engagement from online tutoring videos, noting that engagement is a complex, multidimensional construct influenced by behavioral, emotional, and cognitive states. By analyzing the CASED dataset, the authors highlight the difficulty posed by high inter‑person variability and subjective annotations. They propose a multimodal framework that fuses implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral cues such as head pose, gaze, facial action units, emotion, and wavelet‑based audio features, integrating them via a Perceiver IO bottleneck and modeling participant personalities with variational posteriors. The system employs evidential regression and spectral‑normalized Gaussian process classification heads to provide uncertainty‑aware predictions, achieving competitive performance on the CASED challenge test set while offering well‑calibrated uncertainty metrics.
By Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
arXiv:2608.30563v1 Announce Type: new
Abstract: Multimodal Emotion Recognition (MER) systems often suffer from missing modalities in real-world scenarios. Existing methods usually generate, align, or...
By Jiaqi Zhang, Zheng Pang, Mengting Li, Yiqi Wang, Guangyuan Dong, Chao Xue, Yusen Wu, Zihao Li, Huy Phan, Sicheng Zhao, Bj\"orn W. Schuller, Jiachen Luo
arXiv:2609.34581v2 Announce Type: replace
Abstract: Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, w...
By Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou
arXiv:2606. 20559v1 Announce Type: cross Abstract: Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action.
By Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le, Srijan Das
arXiv:2609.36563v1 Announce Type: new
Abstract: Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible...
By Yihao Qian, Runhao Zeng, Sicheng Zhao, Feng Liang, Hongmin Cai, Mingkui Tan
arXiv:2607. 25961v1 Announce Type: cross Abstract: Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change.
By Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy
The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A...
UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.
By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez