arXiv AI By Vrushank Ahire, Yogesh Kumar, Anouck Girard, M. A. Ganaie

MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening

Read the original on arXiv AI →

MINT (Multimodal Imaging-to-Speech Knowledge Transfer) is a three-stage framework that transfers MRI-derived biomarkers to speech representations for early Alzheimer’s screening. An MRI teacher creates a compact embedding space for CN‑versus‑MCI classification, and a residual projection head aligns speech features to this space using a geometric loss, allowing imaging‑free inference. Experiments on ADNI‑4 show that aligned speech matches speech baselines, while multimodal fusion outperforms MRI alone, and ablations highlight dropout regularization and self‑supervised pretraining as key design choices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 14

A Multimodal Explainable Deep Learning Framework for Alzheimer's Disease Diagnosis using 3D Magnetic Resonance Imaging and Clinical Data

The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.

By Yusuf Brima, Marcellin Atemkeng, Lakshmana Rao Namamula, Antoine Vacavant