From Neurons to Conversation: Speech Brain-Computer Interfaces
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces Open‑Vocabulary Mutual Information (OVMI), an information‑theoretic metric that quantifies how much of a user’s intended speech a speech brain‑computer interface (BCI) can convey relative to a reference word distribution. OVMI enables comparison of systems that use different vocabularies, recording methods, and datasets, revealing that conventional metrics like accuracy and word error rate can overstate performance. Using OVMI, the authors compare existing speech BCI systems, expose trade‑offs between vocabulary coverage and decoding accuracy, and show that optimizing vocabulary selection for OVMI can improve accuracy by up to 16.3% across three speech domains.
The article surveys brain‑to‑language decoding, covering tasks, neural signals, methods, and evaluation across invasive and non‑invasive modalities. It traces the field’s evolution from constrained recognition to text generation, streaming speech, and facial animation, linking tasks to neural populations and decoder representations. The review highlights complementary decoding targets, shared representations, and the growing importance of calibration, feedback, and user control for online communication, while proposing a five‑level future trajectory toward bidirectional cognitive exchange.
This scoping review examines how brain‑computer interfaces (BCIs) can restore sensory and motor functions in people with severe neurological impairment. It introduces a unified 2×2 framework (invasiveness × signal direction) to map 31 key studies, highlighting that most evidence comes from invasive, efferent‑restoration approaches post‑2015. The paper outlines a roadmap toward closed‑loop, bidirectional restoration and identifies gaps in metric standardization, longitudinal data, and cross‑community collaboration.
The 2026 PNPL Competition builds on the 2025 PNPL effort by expanding the LibriBrain dataset to 32 new subjects and more within‑subject data, creating LibriBrain100. It introduces two tracks: a Deep track for high‑performance within‑subject word classification and a Broad track that tests cross‑subject generalisation with progressively less subject‑specific fine‑tuning data, down to 10 minutes. The competition aims to advance non‑invasive brain‑computer interfaces toward practical, clinically feasible communication restoration for people with profound paralysis.
LibriBrain100 is a new large‑scale MEG dataset for speech decoding that contains over 100 hours of high‑quality recordings while subjects listened to naturalistic continuous speech. The dataset more than doubles the size of the original LibriBrain release, with a record 80 hours from a single subject and additional 40‑minute recordings from 32 subjects. The authors demonstrate the value of deep within‑subject data and broad multi‑subject data by achieving state‑of‑the‑art word‑classification performance and showing that supervised fine‑tuning can compensate for limited per‑subject data, all supported by open‑source tools and a public competition leaderboard.
arXiv:2608. 13576v1 Announce Type: cross Abstract: Brain-computer interface (BCI) research relies on multistage computational pipelines, yet progress remains constrained by fragmented data formats, heterogeneous decoder implementations and hardware-specific deployment toolchains, and researchers lack an integrated workflow.