arXiv AI

Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

Cephalonauts One is a 3 Tesla fMRI dataset featuring 30 hours of brain activity recorded from three healthy subjects while they listened to native-language audio podcasts. The release includes the raw fMRI data, corresponding podcast audio, transcript annotations, and derived stimulus embeddings, making it the deepest naturalistic speech fMRI dataset available. A brain‑decoding benchmark is introduced, framing audio segment retrieval as a task where decoders must match fMRI activity to the correct time‑aligned podcast segment, with standardized splits, metrics, and baseline models provided.

arXiv AI
Sep 28

BAT-CLIP: Trimodal Alignment of Brain, Audio and Text

BAT-CLIP is a trimodal alignment framework that jointly aligns intracranial EEG (iEEG) neural embeddings to both pretrained audio and text anchors within a shared, frozen audio‑text manifold. Unlike existing CLIP‑style brain‑speech models that anchor neural activity to a single modality, BAT‑CLIP leverages both audio and text to preserve temporal structure and linguistic separability. On the naturalistic Podcast benchmark, BAT‑CLIP produces more robust representations than bimodal CLIP baselines and demonstrates the value of self‑supervised foundation models for CLIP training.

By Suhyun Kim, Jinmo Han, Danny Dongyeop Han, Ahhyun Lucy Lee, Jewoon Lee, Yonghyeon Gwon, Zach Paris, Chun Kee Chung, Saewoong Bahk, Nam Soo Kim, Seong Jae Hwang, Jiook Cha
arXiv Machine Learning
Aug 27

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

LibriBrain100 is a new large‑scale MEG dataset for speech decoding that contains over 100 hours of high‑quality recordings while subjects listened to naturalistic continuous speech. The dataset more than doubles the size of the original LibriBrain release, with a record 80 hours from a single subject and additional 40‑minute recordings from 32 subjects. The authors demonstrate the value of deep within‑subject data and broad multi‑subject data by achieving state‑of‑the‑art word‑classification performance and showing that supervised fine‑tuning can compensate for limited per‑subject data, all supported by open‑source tools and a public competition leaderboard.

By Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran \"Ozdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones
arXiv AI
Sep 28

Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings

The paper introduces the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which fuses fMRI and MEG data to decode perceived speech from non‑invasive brain signals. Comprehensive experiments show that SICMD improves Top‑1, Top‑10, and Rankacc scores by over 10%, 10%, and 1.7% respectively, while cutting training costs by 88.8% and 60.5% compared to existing multi‑subject and intra‑subject approaches. Visualizations further confirm the method’s effectiveness.

By Aoke Zhang, Jing Chen
arXiv AI
Oct 1

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.

By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu
arXiv AI
Jun 2

MOSS-Audio Technical Report

arXiv:2606. 01802v1 Announce Type: cross Abstract: MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning.

By Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu, Jingqi Chen, Ke Chen, Wenxuan Wang, Yang Wang, Yaozhou Jiang, Yi Jiang, Zhengyuan Lin, Ziqi Chen, Zhaoye Fei, Chenghao Liu, Jun Zhan, Kang Yu, Kexin Huang, Mingshu Chen, Qinyuan Cheng, Ruixiao Li, Shimin Li, Songlin Wang, Yang Gao, Yiyang Zhang, Xipeng Qiu