arXiv AI

MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment

arXiv:2608. 15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion.

arXiv Computer Vision
Sep 4

Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.

By Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel, Mustafa Tevfik Lafci, David Przewozny, Anna Hilsmann, Peter Eisert, Sebastian Bosse
arXiv Machine Learning
5d ago

Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition

The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.

By Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
arXiv AI
Jun 2

UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment

arXiv:2606. 00170v1 Announce Type: cross Abstract: In recent years, emotion recognition based on physiological signals such as electroencephalogram (EEG) has gained considerable attention, as internal physiological data offer greater objectivity and reliability compared to external behavioral data like facial expressions.

By Zheng Wang, Shuo Wang, Junhong Wang
arXiv Computation and Language
Sep 18

Modality Discrepancy Transformer for Ambivalence and Hesitancy Recognition

The paper introduces the Modality Discrepancy Transformer (MDT), a model designed to detect ambivalence and hesitancy in clinical videos by capturing cross‑modal disagreement across facial, vocal, and linguistic signals. MDT expands a 6‑token representation to 9 tokens that include modality embeddings, absolute‑difference features, and Hadamard‑product discrepancy features, which are processed through Transformer self‑attention with FiLM‑based text conditioning and LoRA fine‑tuning. On the BAH dataset from the 3rd ABAW Challenge, MDT achieves a Macro F1 score of 0.7408 on the labelled test split and 0.7368 on the private leaderboard, surpassing the strongest baseline by over 10 points while training in under 20 minutes on a single GPU.

By Shiyu Luo, Yu Wang, Jiawen Huang, Zhaoxiang Xiao, Chenxi Huang, Qi Zhang, Bin Liu
arXiv AI
Aug 12

FUSE: Frame-Unified Stress Estimation from Facial Video

arXiv:2608. 10442v1 Announce Type: cross Abstract: Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose full recordings into short temporal windows before classification.

By Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis
Hugging Face Trending Papers
Jul 23

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video.

arXiv AI
Sep 25

UNWIND: Any-Length Facial Video for Stress Detection without Temporal Windowing

UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.

By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez
arXiv Machine Learning
Sep 17

iMINDBench: iEEG Multi-Institution Neural Decoding Benchmark

iMINDBench is a new benchmark for intracranial electroencephalography (iEEG) neural decoding that evaluates models on fifteen tasks across three naturalistic movie‑watching datasets from multiple institutions. It standardizes preprocessing tracks and evaluation splits to enable consistent comparisons. The study shows that pretrained systems outperform baselines within their tracks, but strong spectral baselines remain competitive, and scaling up supervised data yields only modest or task‑dependent gains.

By Geeling Chau, Saba Hashemi, Yonghyeon Gwon, Eshani Patel, Jan DeWitt, Christopher Wang, Andrii Zahorodnii, Sabera J Talukder, Danny Dongyeop Han, Chun Kee Chung, Maryam M Shanechi, Yisong Yue