arXiv AI

Decoding Error-Related Potentials under Multisensory Feedback with Varying Congruency

arXiv:2607. 24806v1 Announce Type: cross Abstract: Error-related potentials (ErrPs) are widely studied neural signatures associated with error processing in human-machine interaction.

arXiv AI
Sep 12

Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

The paper introduces CoMA-DiT, a bidirectional cross‑modal Diffusion Transformer that uses paired modalities as mutual generative supervision for latent augmentation rather than just inputs for fusion. By conditioning velocity prediction on the paired modality through cross‑modal attention and injecting variation via a reliability‑gated residual mechanism, CoMA‑DiT improves multimodal brain state decoding. Experiments on auditory attention decoding and emotion recognition show consistent gains over 20 baselines, with absolute accuracy and macro‑F1 improvements of 4.28% and 6.70% respectively, and extensive analyses confirm its robustness and interpretability.

By Ziwei Wang, Xingyi He, Hongbin Wang, Tianwang Jia, Bohan Fang, Dongrui Wu
arXiv AI
Sep 10

BIFTA: Brain-Inspired Few-Shot Tactile Adaptation for Unknown Sensors

The paper introduces BIFTA, a Brain‑Inspired Few‑Shot Tactile Adaptation framework that enables a frozen encoder to adapt quickly to an unknown tactile sensor using only a small labeled support set. It preserves pretrained representations via dual‑view statistical memory, builds support‑conditioned spectral graphs to correct sensor‑dependent feature neighborhoods, and employs uncertainty‑gated recurrent propagation to reinforce reliable cross‑query evidence. Benchmarks on three tactile datasets demonstrate that BIFTA dramatically improves adaptation performance, achieving an 87.09% mean Sparsh accuracy on SITR with just 10% labeled data—an increase of 47.22 percentage points over the best prior method.

By Boheng Liu, Ziyu Li, Xia Wu
arXiv AI
Sep 12

RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.

By Xingyi He, Ziwei Wang, Dongrui Wu
arXiv Computer Vision
Aug 24

VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation

VT-MUSE is a multimodal unified sequential representation learning framework for visuotactile manipulation. It uses a two‑stage approach: first, modality‑specific encoders are jointly adapted with cross‑modal temporal alignment and masked‑view consistency; second, a conditional variational latent model processes masked visual sequences and full tactile histories, with auxiliary decoders reconstructing recent visual observations and predicting tactile depth changes. The resulting representation is fed into a lightweight Transformer policy via gated cross‑attention, achieving an 11‑percentage‑point improvement over the strongest baseline in simulation and significant gains in real‑world experiments.

By Congsheng Xu, Qiaochu Yang, Fangyuan Shi, Yifan Han, Baijun Chen, Yiming Wang, Haonan Zhao, Daolin Ma, Xiaokang Yang, Hesheng Wang
Hugging Face Trending Papers
Sep 24

Personalised federated learning for Riemannian and Euclidean EEG decoding

The paper explores personalised federated learning for EEG decoding using two lightweight models: the Riemannian SPDNet and the Euclidean EEGNet. In the personalised approach, all subjects share a common trunk while keeping individual heads, which improves accuracy and reduces communication compared to standard federated learning and centralised training. Experiments on three motor‑imagery datasets show that personalised SPDNet outperforms both standard FL and centralised training, and beats EEGNet on two datasets, though centralised EEGNet remains superior to centralised SPDNet.