The paper presents a method for zero‑shot cross‑subject continuous affect regression using synchronized EEG‑fNIRS data. It decomposes affect trajectories into a shared component across subjects and an individual component derived from a label‑free alpha‑band cross‑channel synchrony marker, which rescales the shared trajectory for each test subject. Extensive validation—including leave‑one‑subject‑out correlation, functional‑form comparison, component ablation, and ceiling analysis—shows the approach achieves lower mean absolute errors than EEGNet and ASAC‑Net baselines on unseen subjects.
By Xuan Wang, Bing Wang, Shuai Chang, Hao Yuan, Xinbo Qi, Xinyue Zhang
arXiv:2608. 15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion.
By Stefanos Gkikas, Eric Nichols, Christian Arzate Cruz, Randy Gomez
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video.
The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.
By Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel, Mustafa Tevfik Lafci, David Przewozny, Anna Hilsmann, Peter Eisert, Sebastian Bosse
The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.
By Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
arXiv:2607. 01400v1 Announce Type: cross Abstract: Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy.
By Barada Sahu, Shivesh Pandey