arXiv:2608. 20186v1 Announce Type: new Abstract: Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable.
By Ingo Marquardt, Anthilia Alchanat, Priyanka Jain
arXiv:2505.12196v2 Announce Type: replace
Abstract: The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conject...
By Yi-Chien Lin, Hongao Zhu, William Schuler
arXiv:2502. 14671v4 Announce Type: replace-cross Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what drives this alignment.
By Maryam Rahimi, Mohammad Reza Daliri, Yadollah Yaghoobzadeh
arXiv:2607. 19394v1 Announce Type: cross Abstract: Generalizing across subjects remains challenging in invasive neural recordings because electrode configurations, anatomical structures, and neural signal patterns vary substantially across individuals.
By Ji-Hoon Heo, Aleksandra Joanna Wisniewska, Seo-Hyun Lee, Seong-Whan Lee
arXiv:2609.24095v1 Announce Type: new
Abstract: While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an u...
By Yueyang Li, Shuran Chen, Wai Ting Siok, Nizhuan Wang
arXiv:2604. 16197v2 Announce Type: replace Abstract: Data attribution and valuation are critical for understanding data-model synergy for Large Language Models (LLMs), yet existing gradient-based methods suffer from scalability challenges on LLMs.
By Yide Ran, Jianwen Xie, Minghui Wang, Wenjin Zheng, Denghui Zhang, Chuan Li, Zhaozhuo Xu
The paper introduces a Cross-Subject Perceived Speech Decoding (CPSD) framework that tackles the challenge of decoding perceived speech from non‑invasive brain recordings across different subjects. CPSD uses a two‑stage training process: first, contrastive learning pre‑trains a source model on multiple subjects to capture shared representations; second, personal specialization fine‑tunes the model for a target subject by extracting consistent components and further training on that subject’s data. A Positional Encoding‑based Spatial Attention (PESA) module is added to remap MEG/EEG data into a standardized reference space, improving cross‑subject consistency. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show that CPSD outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top‑10 accuracy, demonstrating its effectiveness, efficiency, and robustness.
By Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
The study examined whether brain-language model alignment reflects shared computational mechanisms or merely stable lexical‑semantic correspondences. Using whole‑brain encoding across Mandarin, English, and French, transformer representations predicted activity in a distributed network that overlapped across languages and remained stable across layers. Contextual embeddings and measures of prediction or compression did not outperform static lexical embeddings, suggesting that alignment is robust but not informative about shared computational processes.
By Ni Yang, Rui He, Philipp Homan, Iris Sommer, Davide Staub, Wolfram Hinzen
arXiv:2606. 05173v1 Announce Type: cross Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are strongly anchored to surface-form token identity rather than deeper semantic structure.
By Aimen Boukhari
The study investigates whether language models tailored to specific cognitive domains better align with corresponding brain systems. By prompting and fine‑tuning large language models into six domain experts—sensory, spatial, numerical, reasoning, social, and abstract—the authors find that each expert’s representations more closely match the brain region associated with its domain than other experts. This domain‑specific alignment holds across multiple base models and fMRI datasets, while overall prediction accuracy remains largely unchanged, indicating that regional alignment can be obscured when summarizing across the brain.
By Zhivar Sourati, Mengxuan Helen Wu, Nona Ghazizadeh, Jonas Kaplan, Morteza Dehghani, Samuel A. Nastase
While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an unresolved challenge. Decoding individual words fr...
arXiv:2608.29034v1 Announce Type: cross
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...
By Zhang Enyan, R. Thomas McCoy