The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.
By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
The study investigates how data volume, model size, and training duration affect the performance of fMRI foundation models. Using over 200 datasets and 10,000 GPU‑hours, the authors find that larger models benefit more from additional data, and that at a fixed compute budget, increasing data yields greater gains than enlarging the model. By selecting optimal combinations of data, size, and duration, they produce models that outperform existing fMRI foundation models on out‑of‑distribution tasks while requiring less pretraining compute.
By Wenhao Ye, Xuanye Pan, Junfeng Xia, Junxiang Zhang, Mo Wang, Quanying Liu
arXiv:2608.23936v1 Announce Type: new
Abstract: We present a dynamical-systems based model for resting-state functional magnetic resonance imaging (rs-fMRI), trained on a dataset of roughly 40K rs-fM...
By Sourav Pal, Viet Luong, Hoseok Lee, Tingting Dan, Guorong Wu, Richard Davidson, Won Hwa Kim, Vikas Singh
arXiv:2609.31204v1 Announce Type: cross
Abstract: Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models ar...
By Mo Wang, Wenhao Ye, Zihan Ning, Jiayu Zuo, Junfeng Xia, Hongkai Wen, Quanying Liu
Rhamba is a region‑aware pretraining framework for resting‑state fMRI that combines anatomically guided masking with hybrid Attention‑Mamba architectures. The study pretrained models on the ABIDE dataset using three masking strategies (Any, Majority, Pure) and evaluated four architectural variants, finding that the Mamba‑Attention (MA) hybrid achieved the best average AUROC on downstream schizophrenia and ADHD classification tasks. Explainable AI via Integrated Gradients highlighted that performance depends on the interaction between masking strategy and architecture rather than a single dominant configuration.
By Pankaj Pandey, Ruthwik Reddy Doodipala, Pratheek Eranki, Carolina Torres-Rojas, Manob Jyoti Saikia, Ranganatha Sitaram
arXiv:2609.37642v1 Announce Type: new
Abstract: Self-supervised pretraining reshaped prediction in language and vision, and brain foundation models (BFMs) inherited its promise. Representations learn...
By Giovanni Marraffini (UNITO), Victoria Shevchenko (UNITO), Carlo Alberto Barbano (UNITO), Demian Wassermann (MIND)
arXiv:2606. 04772v1 Announce Type: cross Abstract: Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience.
By Hoang-Son Vo, Van-Hung Bui, Minh-Huy Mai-Duc, Tien-Dung Mai, Soo-Hyung Kim
arXiv:2606. 06345v1 Announce Type: cross Abstract: Brain decoding is limited by the availability of labeled neural data, and remains challenging in low-data regimes.
By Yohann Benchetrit, Marl\`ene Careil, Simon Dahan, Hubert Banville, St\'ephane d'Ascoli, Jean-R\'emi King
arXiv:2604. 04958v3 Announce Type: replace-cross Abstract: Recent work suggests that large-scale, multi-animal modeling can significantly improve neural recording analysis.
By Xinhong Xu, Yimeng Zhang, Qichen Qian, Yuanlong Zhang
arXiv:2607. 17782v1 Announce Type: cross Abstract: Foundation models pretrained using self-supervised learning have transformed computer vision by learning transferable representations from large-scale unlabeled data.
By Moona Mazher, Abdul Qayyum, Steven A. Niederer, Daniel C. Alexander
Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question.
arXiv:2609.22271v1 Announce Type: new
Abstract: Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes...
By Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Stephanie Mangesius, Constantin Eisenschink, Philipp Deisl, Michael Knoflach, Astrid E. Grams, Elke R. Gizewski, Rainer Schubert