arXiv AI

Interpretable Alzheimer's Diagnosis via Multimodal Fusion of Regional Brain Experts

arXiv:2512. 10966v3 Announce Type: replace-cross Abstract: Accurate and early diagnosis of Alzheimer's disease (AD) is critical for effective intervention and requires integrating complementary information from multimodal neuroimaging data.

arXiv Computer Vision
Sep 14

A Multimodal Explainable Deep Learning Framework for Alzheimer's Disease Diagnosis using 3D Magnetic Resonance Imaging and Clinical Data

The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.

By Yusuf Brima, Marcellin Atemkeng, Lakshmana Rao Namamula, Antoine Vacavant
arXiv AI
Sep 18

A Two-Stage Multi-Modal MRI Framework for Lifespan Brain Age Prediction

The paper introduces a two-stage multi‑modal MRI framework that processes different MRI modalities independently before integrating them through late fusion. First, the model estimates a probability distribution over six developmental stages, then it predicts age using probability‑weighted, stage‑specialized experts. Experiments across nine datasets—from fetal to elderly—show the method outperforms existing baselines, reducing mean absolute error by 13% and 78% in in‑domain and out‑of‑domain settings, and multi‑modal integration yields 12‑13% performance gains. Analysis on ADNI clinical groups indicates that the predicted brain age gap could help characterize Alzheimer’s‑related brain aging.

By Dingyi Zhang, Ruiying Liu, Yun Wang
arXiv Machine Learning
Jul 31

TIER-MoE: Trust-Informed Expert Routing via Conditional Modality Risk for Multimodal Fusion in Biomedical Classification

arXiv:2607. 27289v1 Announce Type: new Abstract: The promise of multimodal fusion lies in combining complementary sources of evidence, yet more evidence does not always yield a better prediction.

By Yu Chang, Anzhe Cheng, Chenwei Wu, Zhuoran Wang, Jiahao Chen, Tamoghna Chattopadhyay, Sophia I. Thomopoulos, Paul M. Thompson, Liyue Shen, Paul Bogdan
arXiv Computer Vision
Sep 14

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.

By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
arXiv AI
Sep 17

Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging

The paper introduces Generalist‑Specialist Mixture‑of‑Experts (GS‑MoE), a two‑branch architecture that combines a cross‑modal generalist model with modality‑specific specialists through domain‑constrained feature fusion. GS‑MoE improves detection of rare pathologies in multimodal medical imaging, achieving significant per‑class F1 gains and outperforming dense and specialist‑only MoE baselines while using about 53% fewer active parameters at inference. The study demonstrates that balancing cross‑modal shared representations with expert routing can enhance performance on low‑prevalence conditions.

By Johannes Kaiser, Florian Braunmiller, Daniel R\"uckert, Georgios Kaissis
arXiv Computer Vision
Aug 31

3D MRI-Based Alzheimer's Disease Classification Using Multi-Modal 3D CNN with Leakage-Aware Subject-Level Evaluation

The paper presents a multimodal 3D convolutional neural network that classifies Alzheimer’s disease using raw OASIS 1 MRI volumes. It fuses structural T1 images with gray matter, white matter, and cerebrospinal fluid probability maps to capture complementary neuroanatomical information. Evaluated with 5‑fold subject‑level cross‑validation, the model achieves a mean accuracy of 72.34 % and an ROC AUC of 0.7781, with GradCAM visualizations highlighting anatomically relevant regions such as the medial temporal lobe and ventricles.

By Md Sifat, Sania Akter, Akif Islam, Md. Ekramul Hamid, Abu Saleh Musa Miah, Najmul Hassan, Md Abdur Rahim, Jungpil Shin