arXiv Computer Vision

Whole-Body MRI Classification via Prompt-Based Clinical Conditioning

Hugging Face Trending Papers
Jul 27

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence.

arXiv AI
Aug 26

Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis

The paper presents a method for generating cardiac magnetic resonance (CMR) images conditioned on patient metadata using a pretrained latent diffusion model. By encoding structured clinical data and slice position as textual prompts and applying Metadata‑Free Classifier‑Free Guidance, Contrastive Batching, and Inverse‑Frequency Sampling, the authors improve the fidelity of synthetic images, achieving a 57% reduction in Fréchet Inception Distance compared to a baseline without these strategies. Evaluation on 59,058 UK Biobank CMR scans shows better distributional realism and subgroup alignment, though disease‑specific conditioning remains challenging.

By Marc Rodr\'iguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra
arXiv AI
Jun 2

CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

arXiv:2606. 00123v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks.

By Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su, Taiping Qu, Jingwei Guo, Nan Zhang, Hui Wang, Zhen Zhou, Kairui Bo, Yan Chen, Yue Ren, Shuai Li, Lei Xu, Henggui Zhang
arXiv Computer Vision
Aug 27

Hierarchical MoE for Multi-Modal ILD Diagnosis

The paper introduces a hierarchical multimodal mixture-of-experts (MoE) model for interstitial lung disease (ILD) classification. It combines a frozen, pre‑trained imaging expert with structured electronic health records (EHR) through a two‑stage gating system: a modality‑level gate weights imaging and EHR predictions, while a sub‑gating module further decomposes the EHR branch into clinically defined feature groups with learned, group‑specific contributions. The approach preserves stable imaging representations, allows input‑dependent clinical weighting, and enhances interpretability across anatomical regions, imaging–EHR utilization, and EHR feature groups, achieving the highest mean AUC (0.8750 ± 0.0443) under strict patient‑level cross‑validation.

By Alec K. Peltekian, Gorkem Durak, Halil Ertugrul Aktas, Carrie Lynn Richardson, Mary Carns, Kathleen Aren, GR Scott Budinger, Anthony J. Esposito, Alexander Misharin, Alok Nidhi Choudhary, Ankit Agrawal, Ulas Bagci
arXiv Computer Vision
Aug 27

PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology

PANDA (Prototype‑Anchored Data Alignment) is a two‑stage framework that enables a primary‑modality model to benefit from auxiliary modalities even when those modalities are only partially paired or absent at inference. In Stage 1, a shared embedding is learned from the paired subset and class prototypes are estimated from the auxiliary data; in Stage 2, the primary encoder is trained on all subjects using cross‑entropy and alignment to the frozen prototypes. PANDA was evaluated on Alzheimer’s MRI and TCGA‑Lung pathology, achieving significant AUC gains and improved survival prediction while requiring no auxiliary inputs during deployment.

By Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro, Stephan Wunderlich, Rose Dawn Bharat, Siming Bayer, Andreas Maier