arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.
By Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel, Mustafa Tevfik Lafci, David Przewozny, Anna Hilsmann, Peter Eisert, Sebastian Bosse
arXiv:2607. 12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition.
By Aleksei Bakin, Andrey V. Savchenko
arXiv:2609.12668v1 Announce Type: new
Abstract: Recent deepfake detection studies increasingly suggest remote photoplethysmography (rPPG) signals as an authenticity cue. However, existing benchmarks...
By Chenxi Yang, Yassine Ouzar, Larbi Boubchir
arXiv:2409. 00240v2 Announce Type: replace-cross Abstract: Automatic facial action unit (AU) recognition is used widely in facial expression analysis.
By Shuangquan Feng, Virginia R. de Sa
The paper introduces Emo-DVS, a large-scale, multimodal dataset combining event camera, audio, and text data for emotion recognition, designed to mitigate privacy concerns associated with RGB cameras. It proposes the Information‑Guided Gated Fusion (IGF) framework, which pre‑trains an event encoder on the dataset’s FAU subset, adaptively gates modalities to reduce noise, and aligns cross‑modal representations via mutual information maximization. Experiments show that IGF outperforms existing methods on this challenging tri‑modal benchmark.
By Jiaqi Chen, Qinfu Xu, Hao Zhuang, Liyuan Pan
arXiv:2607. 08573v1 Announce Type: new Abstract: Multimodal emotion and sentiment recognition is commonly addressed by early fusion, which concatenates modalities before classification, or late fusion, which combines independently trained unimodal predictors.
By Adis Alihodzic, Selma Skopljakovic Hubljar
arXiv:2606. 11930v1 Announce Type: cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging multimodal learning problem because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.
By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
The paper introduces the Modality Discrepancy Transformer (MDT), a model designed to detect ambivalence and hesitancy in clinical videos by capturing cross‑modal disagreement across facial, vocal, and linguistic signals. MDT expands a 6‑token representation to 9 tokens that include modality embeddings, absolute‑difference features, and Hadamard‑product discrepancy features, which are processed through Transformer self‑attention with FiLM‑based text conditioning and LoRA fine‑tuning. On the BAH dataset from the 3rd ABAW Challenge, MDT achieves a Macro F1 score of 0.7408 on the labelled test split and 0.7368 on the private leaderboard, surpassing the strongest baseline by over 10 points while training in under 20 minutes on a single GPU.
By Shiyu Luo, Yu Wang, Jiawen Huang, Zhaoxiang Xiao, Chenxi Huang, Qi Zhang, Bin Liu
arXiv:2609.09924v1 Announce Type: new
Abstract: Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational conte...
By Oriol Mar\'in, Roger Mar\'i, Gloria Haro, Rafael Redondo
Traits Run Deeper introduces a personality assessment framework that tailors multimodal fusion to each trait dimension. It comprises a Multimodal Foundation Representation module that uses psychology-informed semantic templates, a Trait-Specific Modality Fusion module that asymmetrically fuses modalities to reduce cross‑modal interference, and a Distribution‑Calibrated Personality Regression module that corrects label imbalance. The approach achieves a ~25% reduction in mean squared error on the AVI Challenge 2026 validation set and wins the Personality Assessment Track.
By Jia Li, Qian Chen, Wei Wang, Xinyu Li, Zhenzhen Hu, Dongsheng Shao, Richang Hong, Meng Wang
arXiv:2606. 00170v1 Announce Type: cross Abstract: In recent years, emotion recognition based on physiological signals such as electroencephalogram (EEG) has gained considerable attention, as internal physiological data offer greater objectivity and reliability compared to external behavioral data like facial expressions.
By Zheng Wang, Shuo Wang, Junhong Wang