arXiv AI

Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis

The paper presents a method for generating cardiac magnetic resonance (CMR) images conditioned on patient metadata using a pretrained latent diffusion model. By encoding structured clinical data and slice position as textual prompts and applying Metadata‑Free Classifier‑Free Guidance, Contrastive Batching, and Inverse‑Frequency Sampling, the authors improve the fidelity of synthetic images, achieving a 57% reduction in Fréchet Inception Distance compared to a baseline without these strategies. Evaluation on 59,058 UK Biobank CMR scans shows better distributional realism and subgroup alignment, though disease‑specific conditioning remains challenging.

arXiv AI
Jun 2

CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

arXiv:2606. 00123v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks.

By Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su, Taiping Qu, Jingwei Guo, Nan Zhang, Hui Wang, Zhen Zhou, Kairui Bo, Yan Chen, Yue Ren, Shuai Li, Lei Xu, Henggui Zhang
arXiv Computer Vision
2d ago

CMRVision: A Foundation Model for Cardiac MR Image Analysis

CMRVision is a cardiac magnetic resonance (CMR) foundation model trained with DINOv3-style self‑supervised learning on 36 million multi‑center, multi‑sequence CMR images. It outperforms prior natural‑image, medical‑image, supervised, and CMR baselines on multi‑task segmentation (cine, LGE, mapping) and cine view classification, achieving Dice scores of 0.940–0.967 for LV and 0.855–0.905 for myocardium, and a zero‑shot Dice of 0.692 on unseen LGE long‑axis views. The model demonstrates robust cross‑view generalization and highest average accuracy (0.906) for cine view classification.

By Athira J. Jacob, Puneet Sharma, Daniel Rueckert
arXiv AI
Jul 7

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training

arXiv:2607. 02998v1 Announce Type: cross Abstract: Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning.

By Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
arXiv AI
Jun 19

Scaling Generative Foundation Models for Chest Radiography with Rectified Flow Transformers

arXiv:2606. 19460v1 Announce Type: cross Abstract: We introduce the first generative foundation model for chest radiograph synthesis trained from scratch at the billion-parameter scale.

By Fabio De Sousa Ribeiro, Emma A. M. Stanley, Charles Jones, Tian Xia, Dominic C. Marshall, Laurent Renard Trich\'e, Christopher V. Cosgriff, Panagiotis Dimitrakopoulos, Sotirios A. Tsaftaris, Ben Glocker
arXiv AI
Jul 8

CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training

arXiv:2607. 02998v2 Announce Type: replace-cross Abstract: Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning.

By Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
arXiv AI
Jul 9

CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation

arXiv:2603. 16551v2 Announce Type: replace-cross Abstract: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups.

By Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier
arXiv Computer Vision
3d ago

The MYOSAIQ Challenge: Myocardial Segmentation with Automated Infarct Quantification

arXiv:2608.29246v1 Announce Type: cross Abstract: Late gadolinium enhancement (LGE) cardiac magnetic resonance (MR) imaging is the modality of choice to assess myocardial infarction (MI) lesions. Now...

By Olivier Bernard, William A. Romero R., Cyprien Bouton, Celia Goujat, Hang Jung Ling, Pierre-Marc Jodoin, Fumin Guo, Calder Sheagren, Graham Wright, Abdul Qayyum, Moona Mazher, Steven A. Niederer, Hairui Wang, Xiaomei Wu, Franz Thaler, Gernot Plank, Martin Urschler, Ricardo M. Rosales, Esther Pueyo, Nicolas Duchateau, Frederic Cervenansky, Patrick Clarysse, Loic Belle, Thomas Bochaton, Nathan Mewton, Magalie Viallon, Pierre Croisille
arXiv Machine Learning
Jul 27

Autoregressive EHR Foundation Models with Multimodal Inputs

arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.

By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal