arXiv:2602. 18400v3 Announce Type: replace-cross Abstract: Missing data problems, such as missing modalities in multi-modal brain MRI and missing slices in cardiac MRI, pose significant challenges in clinical practice.
By Junkai Liu, Nay Aung, Theodoros N. Arvanitis, Joao A. C. Lima, Steffen E. Petersen, Le Zhang
arXiv:2607. 06633v1 Announce Type: cross Abstract: In this paper, we address the problem of multimodal federated learning with missing modality.
By Aavash Chhetri, Bibek Niroula, Eduard Vazquez, Yash Raj Shrestha, Prashnna Gyawali, Loris Bazzani, Binod Bhattarai
arXiv:2606. 16484v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and interpreted in natural language.
By Zhiyun Song, Che Liu, Tian Xia, Avinash Kori, Wenjia Bai
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence.
The paper presents a method for generating cardiac magnetic resonance (CMR) images conditioned on patient metadata using a pretrained latent diffusion model. By encoding structured clinical data and slice position as textual prompts and applying Metadata‑Free Classifier‑Free Guidance, Contrastive Batching, and Inverse‑Frequency Sampling, the authors improve the fidelity of synthetic images, achieving a 57% reduction in Fréchet Inception Distance compared to a baseline without these strategies. Evaluation on 59,058 UK Biobank CMR scans shows better distributional realism and subgroup alignment, though disease‑specific conditioning remains challenging.
By Marc Rodr\'iguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra
arXiv:2606. 11107v1 Announce Type: cross Abstract: Clinicians diagnose brain tumors by synthesizing patient symptoms, medical history, and quantitative imaging data from modalities such as MRI and CT scans into a unified clinical judgement.
By Wajih ul Islam, Muhammad Yaqoob, Javed Ali Khan, Volker Steuber
arXiv:2606. 00123v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks.
By Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su, Taiping Qu, Jingwei Guo, Nan Zhang, Hui Wang, Zhen Zhou, Kairui Bo, Yan Chen, Yue Ren, Shuai Li, Lei Xu, Henggui Zhang
The paper introduces a hierarchical multimodal mixture-of-experts (MoE) model for interstitial lung disease (ILD) classification. It combines a frozen, pre‑trained imaging expert with structured electronic health records (EHR) through a two‑stage gating system: a modality‑level gate weights imaging and EHR predictions, while a sub‑gating module further decomposes the EHR branch into clinically defined feature groups with learned, group‑specific contributions. The approach preserves stable imaging representations, allows input‑dependent clinical weighting, and enhances interpretability across anatomical regions, imaging–EHR utilization, and EHR feature groups, achieving the highest mean AUC (0.8750 ± 0.0443) under strict patient‑level cross‑validation.
By Alec K. Peltekian, Gorkem Durak, Halil Ertugrul Aktas, Carrie Lynn Richardson, Mary Carns, Kathleen Aren, GR Scott Budinger, Anthony J. Esposito, Alexander Misharin, Alok Nidhi Choudhary, Ankit Agrawal, Ulas Bagci
arXiv:2512. 10966v3 Announce Type: replace-cross Abstract: Accurate and early diagnosis of Alzheimer's disease (AD) is critical for effective intervention and requires integrating complementary information from multimodal neuroimaging data.
By Farica Zhuang, Shu Yang, Dinara Aliyeva, Zixuan Wen, Duy Duong-Tran, Christos Davatzikos, Tianlong Chen, Song Wang, Li Shen
PANDA (Prototype‑Anchored Data Alignment) is a two‑stage framework that enables a primary‑modality model to benefit from auxiliary modalities even when those modalities are only partially paired or absent at inference. In Stage 1, a shared embedding is learned from the paired subset and class prototypes are estimated from the auxiliary data; in Stage 2, the primary encoder is trained on all subjects using cross‑entropy and alignment to the frozen prototypes. PANDA was evaluated on Alzheimer’s MRI and TCGA‑Lung pathology, achieving significant AUC gains and improved survival prediction while requiring no auxiliary inputs during deployment.
By Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro, Stephan Wunderlich, Rose Dawn Bharat, Siming Bayer, Andreas Maier
arXiv:2605. 23995v4 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data.
By Chathura Wimalasiri, Kishor Nandakishor, Marimuthu Palaniswami