The paper presents Test-Time Adaptation via Cache Personalization (TTA‑CaP), a gradient‑free, cache‑based method that personalizes vision‑language models for facial expression recognition in videos. TTA‑CaP uses three complementary caches—a personalized static cache, a positive target cache, and a negative target cache—controlled by a tri‑gate mechanism to prevent corruption and provide robust subject‑matched evidence. Experiments on BioVid, StressID, and BAH datasets show that TTA‑CaP outperforms state‑of‑the‑art test‑time adaptation methods while keeping computational and memory overhead low.
By Masoumeh Sharafi, Muhammad Osama Zeeshan, Soufiane Belharbi, Alessandro Lameiras Koerich, Marco Pedersoli, Eric Granger
arXiv:2607. 27260v1 Announce Type: new Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge.
By Zhen Zhang, Jielei Chu, Bin Liu, Tianrui Li
arXiv:2608. 06023v1 Announce Type: new Abstract: To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD.
By Bojing Hou, Ruohao Li, Yitong Zhu, Hongjun Liu, Luwen Yu, Yuyang Wang
arXiv:2607. 08839v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference.
By Dominick Reilly, Qiyu Wu, Hiromi Wakaki, Srijan Das, Yuki Mistufuji
arXiv:2608.30563v1 Announce Type: new
Abstract: Multimodal Emotion Recognition (MER) systems often suffer from missing modalities in real-world scenarios. Existing methods usually generate, align, or...
By Jiaqi Zhang, Zheng Pang, Mengting Li, Yiqi Wang, Guangyuan Dong, Chao Xue, Yusen Wu, Zihao Li, Huy Phan, Sicheng Zhao, Bj\"orn W. Schuller, Jiachen Luo
The paper introduces MASA, a method for Wild Test-Time Adaptation that uses a frozen multimodal large language model to provide structured semantic anchors, thereby avoiding the self-referential loop common in existing WTTA techniques. MASA selects a small, diverse set of reliability-ranked anchors, encodes their descriptions, propagates them to nearby test samples, and stores this visual‑semantic information in an online prototype memory. The stored descriptors enable lightweight adaptation of normalization parameters, and MASA is evaluated on the WTTA ImageNet‑C benchmark with ResNet and ViT backbones under limited‑batch, mixed‑domain, and imbalanced‑label‑shift scenarios.
By Zhenbin Wang, Lei Zhang, Lituan Wang, Yan Wang, Zhao Zhang, Wei Huang