arXiv:2608. 06023v1 Announce Type: new Abstract: To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD.
By Bojing Hou, Ruohao Li, Yitong Zhu, Hongjun Liu, Luwen Yu, Yuyang Wang
The paper introduces TTSD‑FAR, a test‑time self‑distillation framework that adapts large video‑language models to missing‑modality scenarios in emotion recognition. A frozen teacher trained on complete modalities guides a low‑rank student, while Fisher‑Anchored Restoration monitors Fisher information to prevent drift and restore the student when distribution shifts occur. Experiments on MELD, DFEW, and BAH with up to 50% missing modalities show TTSD‑FAR consistently outperforms entropy‑based adaptation, retrieval‑augmented generation, and perplexity‑based generation, maintaining performance over long adaptation horizons.
By Muhammad Haseeb Aslam, Alessandro Koerich, Marco Pedersoli, Ali Etemad, Eric Granger
arXiv:2604. 15336v2 Announce Type: replace-cross Abstract: Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone.
By Shuangquan Feng, Laura Fleig, Ruisen Tu, Philip Chi, Edmund Bu, Melinda Ozel, Junhua Ma, Teng Fei, Virginia R. de Sa
arXiv:2608. 14675v1 Announce Type: cross Abstract: While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as photoplethysmography (PPG), its suitability for highly subjective tasks remains unproven.
By Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski
The paper tackles the challenge of predicting student engagement from online tutoring videos, noting that engagement is a complex, multidimensional construct influenced by behavioral, emotional, and cognitive states. By analyzing the CASED dataset, the authors highlight the difficulty posed by high inter‑person variability and subjective annotations. They propose a multimodal framework that fuses implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral cues such as head pose, gaze, facial action units, emotion, and wavelet‑based audio features, integrating them via a Perceiver IO bottleneck and modeling participant personalities with variational posteriors. The system employs evidential regression and spectral‑normalized Gaussian process classification heads to provide uncertainty‑aware predictions, achieving competitive performance on the CASED challenge test set while offering well‑calibrated uncertainty metrics.
By Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
arXiv:2608.30563v1 Announce Type: new
Abstract: Multimodal Emotion Recognition (MER) systems often suffer from missing modalities in real-world scenarios. Existing methods usually generate, align, or...
By Jiaqi Zhang, Zheng Pang, Mengting Li, Yiqi Wang, Guangyuan Dong, Chao Xue, Yusen Wu, Zihao Li, Huy Phan, Sicheng Zhao, Bj\"orn W. Schuller, Jiachen Luo
arXiv:2606. 20559v1 Announce Type: cross Abstract: Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action.
By Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le, Srijan Das
The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A...
arXiv:2609.34581v2 Announce Type: replace
Abstract: Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, w...
By Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou
Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings.
arXiv:2609.13024v1 Announce Type: new
Abstract: As a key model compression technique, knowledge distillation aims to transfer knowledge from a high-capacity teacher model to a lightweight student mod...
By Yanjiang Shi, Peng Zhao, Nan Qi, Guiqin Wang
arXiv:2609.36563v1 Announce Type: new
Abstract: Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible...
By Yihao Qian, Runhao Zeng, Sicheng Zhao, Feng Liang, Hongmin Cai, Mingkui Tan