Evaluating the Generalization of Neuroimaging Foundation Models on African Brain MRI
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
BrainFedFM is a structural brain MRI foundation model that was federatively pretrained on 164,707 3‑D scans from 42 sites using a dual‑priority approach that emphasizes informative anatomical regions locally and prioritizes site contributions globally. The model outperformed seven baseline models—including four centralized foundation models—across 20 downstream tasks (classification, regression, segmentation), achieving a mean rank of 1.68 and a 50% performance gain, especially in classification and regression and among underrepresented populations. These results demonstrate the model’s generalizability and show that federated pretraining can effectively develop neuroimaging foundation models without pooling raw images.
arXiv:2603. 28387v2 Announce Type: replace Abstract: Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts.
The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.
The study evaluates whether disease can be identified from reactive, non‑lesional brain tissue in intracranial biopsies. Using four foundation‑model encoders within an attention‑based multiple‑instance learning framework on 245 whole‑slide images, the authors find that disease labels remain predictive even after controlling for slide size and sampling bias, and that performance is similar across all encoders. Signed instance‑contribution maps and expert review confirm that predictive signals localize to reactive parenchyma rather than artifacts such as blood. "whyItMatters":"The findings demonstrate that weakly supervised models can recover disease signals from tissue traditionally considered non‑diagnostic, highlighting the need for provenance‑only baselines in computational pathology benchmarks."
The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.
arXiv:2603.28387v3 Announce Type: replace-cross Abstract: Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language model...