arXiv Computer Vision

Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction

arXiv Computer Vision
Sep 14

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.

By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
arXiv Machine Learning
Sep 24

A Scaling Study for fMRI Foundation Models

The study investigates how data volume, model size, and training duration affect the performance of fMRI foundation models. Using over 200 datasets and 10,000 GPU‑hours, the authors find that larger models benefit more from additional data, and that at a fixed compute budget, increasing data yields greater gains than enlarging the model. By selecting optimal combinations of data, size, and duration, they produce models that outperform existing fMRI foundation models on out‑of‑distribution tasks while requiring less pretraining compute.

By Wenhao Ye, Xuanye Pan, Junfeng Xia, Junxiang Zhang, Mo Wang, Quanying Liu
arXiv AI
Sep 21

Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI

Rhamba is a region‑aware pretraining framework for resting‑state fMRI that combines anatomically guided masking with hybrid Attention‑Mamba architectures. The study pretrained models on the ABIDE dataset using three masking strategies (Any, Majority, Pure) and evaluated four architectural variants, finding that the Mamba‑Attention (MA) hybrid achieved the best average AUROC on downstream schizophrenia and ADHD classification tasks. Explainable AI via Integrated Gradients highlighted that performance depends on the interaction between masking strategy and architecture rather than a single dominant configuration.

By Pankaj Pandey, Ruthwik Reddy Doodipala, Pratheek Eranki, Carolina Torres-Rojas, Manob Jyoti Saikia, Ranganatha Sitaram
Hugging Face Trending Papers
Jun 3

Coarse-to-fine Hierarchical Architecture with Sequential Mamba for Brain Reconstruction

Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question.

arXiv Computer Vision
Sep 22

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

arXiv:2609.22271v1 Announce Type: new Abstract: Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes...

By Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Stephanie Mangesius, Constantin Eisenschink, Philipp Deisl, Michael Knoflach, Astrid E. Grams, Elke R. Gizewski, Rainer Schubert