The study investigates how data volume, model size, and training duration affect the performance of fMRI foundation models. Using over 200 datasets and 10,000 GPU‑hours, the authors find that larger models benefit more from additional data, and that at a fixed compute budget, increasing data yields greater gains than enlarging the model. By selecting optimal combinations of data, size, and duration, they produce models that outperform existing fMRI foundation models on out‑of‑distribution tasks while requiring less pretraining compute.
By Wenhao Ye, Xuanye Pan, Junfeng Xia, Junxiang Zhang, Mo Wang, Quanying Liu
arXiv:2609.34167v2 Announce Type: replace
Abstract: Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost....
By Juhyeon Park, Yeonwoo Kim, Peter Yongho Kim, Yansen Wang, Mingqing Xiao, Dongqi Han, Dongsheng Li, Taesup Moon
arXiv:2606. 04772v1 Announce Type: cross Abstract: Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience.
By Hoang-Son Vo, Van-Hung Bui, Minh-Huy Mai-Duc, Tien-Dung Mai, Soo-Hyung Kim
The paper presents a reproducible single‑subject baseline for reconstructing visual stimuli from EEG using a temporal‑spatial convolutional encoder that maps averaged EEG signals to 512‑dimensional ViT-B/32 image features. On the THINGS‑EEG2 dataset, the model achieves 12.83%, 39.17%, and 58.00% image recall at ranks 1, 5, and 10, respectively, outperforming analytical chance levels. The study also shows that performance drops sharply when applying a model trained on one subject to others, and that direct conditional generators without external visual weights produce noise‑dominated outputs, indicating that only coarse semantic decoding is feasible under the tested protocol.
By Harshit Goyal
arXiv:2606. 00121v1 Announce Type: cross Abstract: Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task in brain decoding.
By Yizhuo Lu, Changde Du, Qiongyi Zhou, Liuyun Jiang, Huiguang He
arXiv:2607. 22753v1 Announce Type: cross Abstract: Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical advances.
By Rishab S. Iyer, Jiaxin Cindy Tu, Cesar Kadir Torrico Villanueva, Anish Mahishi, Ross P. Kempner, Jacob S. Prince, Ernest W. Lo, Akash Bhowmick, Hritik Arasu, Amaar Chughtai, Elizabeth A. McDevitt, Paul S. Scotti, Kenneth A. Norman
arXiv:2603.02767v4 Announce Type: replace-cross
Abstract: Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield repre...
By Hanpeng Liu, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, Kaiwen Long, Yaqian Li, Kun He
arXiv:2609.36366v1 Announce Type: cross
Abstract: Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used d...
By Iishaan Inabathini, Margaret M. Henderson
Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question.
arXiv:2503. 22697v3 Announce Type: replace-cross Abstract: Decoding sensory experiences from neural activity to reconstruct human-perceived visual stimuli and semantic content remains a challenge in neuroscience and artificial intelligence.
By Feihan Feng, Jingxin Nie
NeuroSketch presents a practical design recipe for neural decoding, beginning with a comparative study of nine basic architectures that identifies CNN‑2D as the most effective. The recipe incorporates macro‑level gradual feature‑map expansion and early downsampling, along with micro‑level grouped convolutions, resulting in two variants—NeuroSketch‑Base (1.4M parameters) and NeuroSketch‑Large (4.2M parameters). Across nearly 5,000 experiments on eight tasks involving visual, auditory, and speech modalities and EEG, SEEG, and ECoG signals, both variants outperform ten baseline models on every task.
By Gaorui Zhang, Zhizhang Yuan, Jialan Yang, Junru Chen, Fanqi Shen, Li Meng, Yang Yang
This survey reviews recent advances in converting non‑invasive EEG signals into images, text, and audio using generative AI techniques such as GANs, VAEs, transformers, and diffusion models. It summarizes datasets, feature‑encoding methods, evaluation metrics, and key challenges, noting that EEG‑to‑image models mainly use encoder‑decoder architectures, EEG‑to‑text leverages transformer language models, and EEG‑to‑audio maps signals to mel‑spectrograms for vocoder synthesis. The paper highlights the limitations of small, heterogeneous datasets, poor cross‑subject generalization, and the lack of standardized benchmarks, while providing open‑source resources to support reproducible research.
By Shreya Shukla, Jose Torres, Akshaj Murhekar, Christina Liu, Abhijit Mishra, Jacek Gwizdka, Shounak Roychowdhury