arXiv AI By Minghao Kong, Jiurun Chen, Ying Gao, Xiangbin Meng, Rongjie Wang

Low-Cost Video--Time Priors as a Strong Baseline for EEG--fNIRS Emotion Regression on Familiar Videos

Read the original on arXiv AI →

The paper presents a low‑cost baseline for continuous emotion regression that relies on video–time priors when users watch familiar videos. It compares this baseline to a fusion of EEG and fNIRS signals, finding that the video–time prior alone achieves mean absolute errors within 0.05 and 0.32 of the fusion model in internal and external evaluations. Ablation studies show that video identity and within‑video time explain most of the performance, while EEG–fNIRS contributions are smaller and variable across participants and videos.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Modeling Shared and Individual Structure for Cross-Subject Continuous Affect Regression from EEG-fNIRS

The paper presents a method for zero‑shot cross‑subject continuous affect regression using synchronized EEG‑fNIRS data. It decomposes affect trajectories into a shared component across subjects and an individual component derived from a label‑free alpha‑band cross‑channel synchrony marker, which rescales the shared trajectory for each test subject. Extensive validation—including leave‑one‑subject‑out correlation, functional‑form comparison, component ablation, and ceiling analysis—shows the approach achieves lower mean absolute errors than EEGNet and ASAC‑Net baselines on unseen subjects.

By Xuan Wang, Bing Wang, Shuai Chang, Hao Yuan, Xinbo Qi, Xinyue Zhang
Hugging Face Trending Papers
Jul 23

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video.

arXiv Computer Vision
Sep 4

Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.

By Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel, Mustafa Tevfik Lafci, David Przewozny, Anna Hilsmann, Peter Eisert, Sebastian Bosse
arXiv Machine Learning
Sep 29

Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition

The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.

By Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen