The paper presents a method for zero‑shot cross‑subject continuous affect regression using synchronized EEG‑fNIRS data. It decomposes affect trajectories into a shared component across subjects and an individual component derived from a label‑free alpha‑band cross‑channel synchrony marker, which rescales the shared trajectory for each test subject. Extensive validation—including leave‑one‑subject‑out correlation, functional‑form comparison, component ablation, and ceiling analysis—shows the approach achieves lower mean absolute errors than EEGNet and ASAC‑Net baselines on unseen subjects.
By Xuan Wang, Bing Wang, Shuai Chang, Hao Yuan, Xinbo Qi, Xinyue Zhang
arXiv:2608. 15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion.
By Stefanos Gkikas, Eric Nichols, Christian Arzate Cruz, Randy Gomez
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video.
The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.
By Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel, Mustafa Tevfik Lafci, David Przewozny, Anna Hilsmann, Peter Eisert, Sebastian Bosse
The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.
By Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
arXiv:2607. 01400v1 Announce Type: cross Abstract: Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy.
By Barada Sahu, Shivesh Pandey
arXiv:2606. 11555v1 Announce Type: cross Abstract: The escalating demand for mental healthcare, driven by rising societal stress, highlights the limitations of traditional psychiatric diagnostics.
By Riki Sakurai, Simon Kojima, Mihoko Otake-Matsuura, Shin'ichiro Kanoh, Tomasz M. Rutkowski
arXiv:2606. 30104v1 Announce Type: new Abstract: Electroencephalography (EEG) foundation models aim to learn generalizable representations from large-scale brain recordings.
By Ay\c{s}e Bet\"ul Y\"uce, Chris Joey Leffler, Sarun Varghese, Myra Spiliopoulou, Sebastian Stober
The paper introduces AFOR, a tensor‑wise adaptive optimizer for EEG decoding that replaces the fixed second‑moment decay coefficient used in Adam/AdamW with a dynamic coefficient estimated online from local gradient state. AFOR combines a Residual‑Alignment Signal Scorer (RASS) to assess gradient quality and an Adaptive Forgetting Controller (AFC) to map this score to a bounded per‑step decay coefficient, with cumulative‑product initialization correction for consistency. In cross‑subject experiments on three EEG benchmarks, AFOR outperforms standard optimizers, improving mean test accuracy by 3.00%, 2.07%, and 4.38% over Adam.
By Hongyu Zhu, Lin Chen, Jing Chen, Yuting Zhou, Mingsheng Shang
iMINDBench is a new benchmark for intracranial electroencephalography (iEEG) neural decoding that evaluates models on fifteen tasks across three naturalistic movie‑watching datasets from multiple institutions. It standardizes preprocessing tracks and evaluation splits to enable consistent comparisons. The study shows that pretrained systems outperform baselines within their tracks, but strong spectral baselines remain competitive, and scaling up supervised data yields only modest or task‑dependent gains.
By Geeling Chau, Saba Hashemi, Yonghyeon Gwon, Eshani Patel, Jan DeWitt, Christopher Wang, Andrii Zahorodnii, Sabera J Talukder, Danny Dongyeop Han, Chun Kee Chung, Maryam M Shanechi, Yisong Yue
arXiv:2608. 06023v1 Announce Type: new Abstract: To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD.
By Bojing Hou, Ruohao Li, Yitong Zhu, Hongjun Liu, Luwen Yu, Yuyang Wang
NeuroAtlas is the largest EEG benchmark to date, comprising 42 datasets and 260,000 hours of clinical EEG data across epilepsy, sleep medicine, brain age estimation, and brain‑computer interfaces. The study evaluates foundation models (FMs) for EEG against supervised baselines and generic time‑series FMs, finding that EEG‑specific FMs do not consistently outperform generic ones. It also demonstrates that standard machine‑learning metrics are inadequate for clinical relevance, advocating for task‑specific measures such as event‑level decision quality, hypnogram features, and brain‑age gap.
By Konstantinos Kontras, Trui Osselaer, Stylianos G. Mouslech, Angeliki-Ilektra Karaiskou, Guido Gagliardi, Thomas Strypsteen, Mohammad Hossein Badiei, Anku Rani, Maarten Vanmarcke, Miguel Bhagubai, Chanakya Ekbote, Jaedong Hwang, Christos Chatzichristos, Paul Pu Liang, Maarten De Vos