arXiv AI

MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening

MINT (Multimodal Imaging-to-Speech Knowledge Transfer) is a three-stage framework that transfers MRI-derived biomarkers to speech representations for early Alzheimer’s screening. An MRI teacher creates a compact embedding space for CN‑versus‑MCI classification, and a residual projection head aligns speech features to this space using a geometric loss, allowing imaging‑free inference. Experiments on ADNI‑4 show that aligned speech matches speech baselines, while multimodal fusion outperforms MRI alone, and ablations highlight dropout regularization and self‑supervised pretraining as key design choices.

arXiv Computer Vision
Sep 14

A Multimodal Explainable Deep Learning Framework for Alzheimer's Disease Diagnosis using 3D Magnetic Resonance Imaging and Clinical Data

The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.

By Yusuf Brima, Marcellin Atemkeng, Lakshmana Rao Namamula, Antoine Vacavant
arXiv AI
6d ago

M$^2$PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease

M$^2$PFN is an end‑to‑end multimodal framework that extends the TabPFN in‑context learning engine to Alzheimer’s disease diagnosis by aligning 3D‑MRI and tabular features in a shared subspace. It performs differentiable inference through TabPFN’s transformer, back‑propagates gradients into the encoders, and incorporates a frozen tabular‑only prediction via a gated shortcut. On the ADNI cohort it achieves 65.55 % macro‑F1 and 82.21 % macro‑AUC, surpassing unimodal and multimodal baselines, and it generalizes to external cohorts without retraining.

By Lujia Zhong, Shuo Huang, Jianwei Zhang, Xinyu Nie, Yonggang Shi
arXiv AI
6d ago

Cross-Task Generalization in Handwriting-Based Alzheimer's Screening via Vision Language Adaptation

The paper introduces a lightweight Cross‑Layer Fusion Adapter (CLFA) that adapts the CLIP vision‑language model for handwriting‑based Alzheimer's disease screening. CLFA inserts multi‑level adapters into a frozen visual encoder, fusing cross‑layer features with depthwise 2D convolutions to capture both local stroke irregularities and higher‑level handwriting structure. On the Darwin dataset, CLFA achieves 74.63% AUC, 74.85% accuracy, and 73.72% F1, outperforming the best competing model by 2.15, 1.79, and 1.87 percentage points across 600 task‑disjoint source‑target pairs.

By Changqing Gong, Huafeng Qin, Mounim A. El-Yacoubi
arXiv Machine Learning
Sep 23

MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction

MMAP is a Multimodal Missing‑Aware Alignment Pretraining method designed to learn image‑tabular representations from incomplete data. It uses a sigmoid contrastive learning image encoder with generative reconstruction, a tabular encoder based on a foundation model, and a missing token generator to handle missing modalities. The approach is evaluated on longitudinal Alzheimer’s tasks—predicting disease stage conversion and amyloid status—and outperforms both multimodal and unimodal baselines.

By Fiona Kekwick, Matthew Baugh, Bernhard Kainz, Paul M. Matthews, Wenjia Bai