arXiv:2606. 16484v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and interpreted in natural language.
By Zhiyun Song, Che Liu, Tian Xia, Avinash Kori, Wenjia Bai
arXiv:2606. 09907v1 Announce Type: cross Abstract: Multimodal clinical learning is increasingly important for integrating diverse patient data, including imaging, text, and personalised health records.
By Maxx Richard Rahman, Prakhar Kumar, Wolfgang Maass
arXiv:2607. 11656v1 Announce Type: cross Abstract: Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data.
By Christelle Schneuwly Diaz, Narmina Baghirova, Duy-Thanh Vu, Duy-Cat Can, Gilles Allali, Philippe Ryvlin, Oliver Y. Ch\'en
M$^2$PFN is an end‑to‑end multimodal framework that extends the TabPFN in‑context learning engine to Alzheimer’s disease diagnosis by aligning 3D‑MRI and tabular features in a shared subspace. It performs differentiable inference through TabPFN’s transformer, back‑propagates gradients into the encoders, and incorporates a frozen tabular‑only prediction via a gated shortcut. On the ADNI cohort it achieves 65.55 % macro‑F1 and 82.21 % macro‑AUC, surpassing unimodal and multimodal baselines, and it generalizes to external cohorts without retraining.
By Lujia Zhong, Shuo Huang, Jianwei Zhang, Xinyu Nie, Yonggang Shi
arXiv:2606. 06328v1 Announce Type: new Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because electrodes detach or an entire respiratory channel is unavailable during overnight monitoring.
By Ziwen Kan, Wugeng Zheng, Tianlong Chen, Song Wang
MINT (Multimodal Imaging-to-Speech Knowledge Transfer) is a three-stage framework that transfers MRI-derived biomarkers to speech representations for early Alzheimer’s screening. An MRI teacher creates a compact embedding space for CN‑versus‑MCI classification, and a residual projection head aligns speech features to this space using a geometric loss, allowing imaging‑free inference. Experiments on ADNI‑4 show that aligned speech matches speech baselines, while multimodal fusion outperforms MRI alone, and ablations highlight dropout regularization and self‑supervised pretraining as key design choices.
By Vrushank Ahire, Yogesh Kumar, Anouck Girard, M. A. Ganaie
arXiv:2608.30824v1 Announce Type: new
Abstract: Combining whole-body magnetic resonance imaging (WB-MRI) with clinical variables has the potential to improve systemic disease diagnosis by leveraging...
By Laura Daza, Marta Hasny, Cristina Gonz\'alez, Julia A. Schnabel
arXiv:2604. 22700v2 Announce Type: replace Abstract: Modeling and predicting neurodegenerative disease progression from medical images remains a major challenge in medical AI, with significant implications for early diagnosis, disease monitoring, and treatment planning.
By Nivetha Jayakumar, Swakshar Deb, Bahram Jafrasteh, Qingyu Zhao, Miaomiao Zhang
arXiv:2606. 11794v1 Announce Type: cross Abstract: Neurodegenerative diseases such as Alzheimer's disease (AD) require accurate and scalable tools for assessing disease severity, yet current clinical staging remains time-intensive and prone to variability.
By Boris-Stephan Rauchmann, Jonathan Laib, Buse Ercik, Robert Perneczky, Sergio Altares-L\'opez
NOAH is a generative transformer that learns the full multimodal patient journey by integrating bidirectional time and a variational latent space. Trained on over 559 million clinical events from 431,000 hospital visits, it processes medical images, time‑series, numeric signals, categorical events, and both structured and unstructured records. The model supports autoregressive forecasting, zero‑shot classification, and counterfactual intervention simulation, yielding strong performance on clinical outcomes, ICD chapters, comorbidities, and time‑to‑event prediction.
By Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov, \"Ozg\"un Turgut, Michelle Espranita Liman, Lisa Steinhelfer, Rickmer Braren, Daniel Rueckert
The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.
By Yusuf Brima, Marcellin Atemkeng, Lakshmana Rao Namamula, Antoine Vacavant
PANDA (Prototype‑Anchored Data Alignment) is a two‑stage framework that enables a primary‑modality model to benefit from auxiliary modalities even when those modalities are only partially paired or absent at inference. In Stage 1, a shared embedding is learned from the paired subset and class prototypes are estimated from the auxiliary data; in Stage 2, the primary encoder is trained on all subjects using cross‑entropy and alignment to the frozen prototypes. PANDA was evaluated on Alzheimer’s MRI and TCGA‑Lung pathology, achieving significant AUC gains and improved survival prediction while requiring no auxiliary inputs during deployment.
By Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro, Stephan Wunderlich, Rose Dawn Bharat, Siming Bayer, Andreas Maier