arXiv Machine Learning

MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction

MMAP is a Multimodal Missing‑Aware Alignment Pretraining method designed to learn image‑tabular representations from incomplete data. It uses a sigmoid contrastive learning image encoder with generative reconstruction, a tabular encoder based on a foundation model, and a missing token generator to handle missing modalities. The approach is evaluated on longitudinal Alzheimer’s tasks—predicting disease stage conversion and amyloid status—and outperforms both multimodal and unimodal baselines.

arXiv Machine Learning
Jul 14

Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohorts

arXiv:2607. 11656v1 Announce Type: cross Abstract: Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data.

By Christelle Schneuwly Diaz, Narmina Baghirova, Duy-Thanh Vu, Duy-Cat Can, Gilles Allali, Philippe Ryvlin, Oliver Y. Ch\'en
arXiv AI
6d ago

M$^2$PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease

M$^2$PFN is an end‑to‑end multimodal framework that extends the TabPFN in‑context learning engine to Alzheimer’s disease diagnosis by aligning 3D‑MRI and tabular features in a shared subspace. It performs differentiable inference through TabPFN’s transformer, back‑propagates gradients into the encoders, and incorporates a frozen tabular‑only prediction via a gated shortcut. On the ADNI cohort it achieves 65.55 % macro‑F1 and 82.21 % macro‑AUC, surpassing unimodal and multimodal baselines, and it generalizes to external cohorts without retraining.

By Lujia Zhong, Shuo Huang, Jianwei Zhang, Xinyu Nie, Yonggang Shi
arXiv AI
Sep 17

MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening

MINT (Multimodal Imaging-to-Speech Knowledge Transfer) is a three-stage framework that transfers MRI-derived biomarkers to speech representations for early Alzheimer’s screening. An MRI teacher creates a compact embedding space for CN‑versus‑MCI classification, and a residual projection head aligns speech features to this space using a geometric loss, allowing imaging‑free inference. Experiments on ADNI‑4 show that aligned speech matches speech baselines, while multimodal fusion outperforms MRI alone, and ablations highlight dropout regularization and self‑supervised pretraining as key design choices.

By Vrushank Ahire, Yogesh Kumar, Anouck Girard, M. A. Ganaie
arXiv Computer Vision
Aug 21

4DLoG: Generative Modeling of Neurodegenerative Brain Anatomy with 4D Longitudinal Diffusion Model

arXiv:2604. 22700v2 Announce Type: replace Abstract: Modeling and predicting neurodegenerative disease progression from medical images remains a major challenge in medical AI, with significant implications for early diagnosis, disease monitoring, and treatment planning.

By Nivetha Jayakumar, Swakshar Deb, Bahram Jafrasteh, Qingyu Zhao, Miaomiao Zhang
arXiv AI
Sep 10

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

NOAH is a generative transformer that learns the full multimodal patient journey by integrating bidirectional time and a variational latent space. Trained on over 559 million clinical events from 431,000 hospital visits, it processes medical images, time‑series, numeric signals, categorical events, and both structured and unstructured records. The model supports autoregressive forecasting, zero‑shot classification, and counterfactual intervention simulation, yielding strong performance on clinical outcomes, ICD chapters, comorbidities, and time‑to‑event prediction.

By Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov, \"Ozg\"un Turgut, Michelle Espranita Liman, Lisa Steinhelfer, Rickmer Braren, Daniel Rueckert
arXiv Computer Vision
Sep 14

A Multimodal Explainable Deep Learning Framework for Alzheimer's Disease Diagnosis using 3D Magnetic Resonance Imaging and Clinical Data

The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.

By Yusuf Brima, Marcellin Atemkeng, Lakshmana Rao Namamula, Antoine Vacavant
arXiv Computer Vision
Aug 27

PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology

PANDA (Prototype‑Anchored Data Alignment) is a two‑stage framework that enables a primary‑modality model to benefit from auxiliary modalities even when those modalities are only partially paired or absent at inference. In Stage 1, a shared embedding is learned from the paired subset and class prototypes are estimated from the auxiliary data; in Stage 2, the primary encoder is trained on all subjects using cross‑entropy and alignment to the frozen prototypes. PANDA was evaluated on Alzheimer’s MRI and TCGA‑Lung pathology, achieving significant AUC gains and improved survival prediction while requiring no auxiliary inputs during deployment.

By Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro, Stephan Wunderlich, Rose Dawn Bharat, Siming Bayer, Andreas Maier