arXiv Computer Vision

A Multimodal Explainable Deep Learning Framework for Alzheimer's Disease Diagnosis using 3D Magnetic Resonance Imaging and Clinical Data

The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.

arXiv Computer Vision
Sep 7

A Generalizable Feature Extractor for Alzheimer's-Related Brain MRI Tasks

The study investigates whether a compact, supervised 3D CNN pretrained for brain‑age prediction can act as a reusable foundation model for various Alzheimer's‑related neuroimaging tasks. By freezing the 7.18 million weights and adding only ~1 % of trainable parameters via Low‑Rank Adaptation, the model achieved high performance across six experiments, including dementia classification, MCI progression prediction, amyloid positivity detection, and volume estimation of hippocampal and white matter hypointensities. The results demonstrate that the pretrained brain‑age model generalizes well to new datasets without retraining, offering a data‑efficient alternative to larger networks.

By Reza Rajabli, D. Louis Collins
arXiv Computer Vision
Aug 31

3D MRI-Based Alzheimer's Disease Classification Using Multi-Modal 3D CNN with Leakage-Aware Subject-Level Evaluation

The paper presents a multimodal 3D convolutional neural network that classifies Alzheimer’s disease using raw OASIS 1 MRI volumes. It fuses structural T1 images with gray matter, white matter, and cerebrospinal fluid probability maps to capture complementary neuroanatomical information. Evaluated with 5‑fold subject‑level cross‑validation, the model achieves a mean accuracy of 72.34 % and an ROC AUC of 0.7781, with GradCAM visualizations highlighting anatomically relevant regions such as the medial temporal lobe and ventricles.

By Md Sifat, Sania Akter, Akif Islam, Md. Ekramul Hamid, Abu Saleh Musa Miah, Najmul Hassan, Md Abdur Rahim, Jungpil Shin
arXiv Machine Learning
Jul 14

Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohorts

arXiv:2607. 11656v1 Announce Type: cross Abstract: Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data.

By Christelle Schneuwly Diaz, Narmina Baghirova, Duy-Thanh Vu, Duy-Cat Can, Gilles Allali, Philippe Ryvlin, Oliver Y. Ch\'en
arXiv Machine Learning
Aug 27

Modality Contribution Score - A Per-Patient Framework for Quantifying the Relative Diagnostic Contribution of Structural MRI and Amyloid PET in Alzheimer's Disease

The paper introduces MCNet, a neural network that assigns a Modality Contribution Score (MCS) to each patient, indicating how much structural MRI versus amyloid PET drives the diagnostic decision for Alzheimer’s disease. Across 327 ADNI-3 participants, MCNet achieved strong three‑class staging (AUC = 0.881) and the MCS showed a clear, statistically significant increase in PET dominance from cognitively normal to AD. The method was validated on an independent OASIS‑3 cohort and compared favorably to SHAP, suggesting it can guide personalized imaging and clinical trial decisions.

By Dawa Chyophel Lepcha, Aaliya Ali, Sophie A. Martin, Deepika Koundal, Pierrick Coupe, Shabbir Syed-Abdul
arXiv Computer Vision
4d ago

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.

By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian