arXiv Computer Vision
Sep 2

CMRVision: A Foundation Model for Cardiac MR Image Analysis

CMRVision is a cardiac magnetic resonance (CMR) foundation model trained with DINOv3-style self‑supervised learning on 36 million multi‑center, multi‑sequence CMR images. It outperforms prior natural‑image, medical‑image, supervised, and CMR baselines on multi‑task segmentation (cine, LGE, mapping) and cine view classification, achieving Dice scores of 0.940–0.967 for LV and 0.855–0.905 for myocardium, and a zero‑shot Dice of 0.692 on unseen LGE long‑axis views. The model demonstrates robust cross‑view generalization and highest average accuracy (0.906) for cine view classification.

By Athira J. Jacob, Puneet Sharma, Daniel Rueckert
arXiv AI
Jun 2

CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

arXiv:2606. 00123v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks.

By Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su, Taiping Qu, Jingwei Guo, Nan Zhang, Hui Wang, Zhen Zhou, Kairui Bo, Yan Chen, Yue Ren, Shuai Li, Lei Xu, Henggui Zhang
arXiv AI
Sep 15

MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management

arXiv:2603.22179v2 Announce Type: replace Abstract: Cardiovascular disease remains the leading cause of global mortality, with progress hindered by human interpretation of complex cardiac tests. Curr...

By Jack W O'Sullivan, Mohammad Asadi, Lennart Elbe, Akshay Chaudhari, Tahoura Nedaee, Francois Haddad, Ivan Lopez, Fang Cao, Michael Salerno, Li Fe-Fei, Ehsan Adeli, Rima Arnaout, Euan A Ashley
arXiv Computer Vision
Sep 14

SV-Cine: Diagnosis-Conditioned Segmentation of Single Ventricle Physiology via Generative Data Augmentation

The paper introduces SV-Cine, a cardiac MRI segmentation framework tailored for single ventricle physiology (SVP). It combines a generative data augmentation pipeline that creates synthetic 3D cardiac meshes and MRI, with a diagnosis-conditioned adaptation of the CineMA foundation model that uses patient-level diagnostic information to improve segmentation. Evaluations on an internal cohort show high Dice scores for left and right ventricles, outperforming nnU-Net, and demonstrate that incorporating diagnosis priors can adapt a pretrained model to specialized SVP tasks.

By Lila Cunge, Yuehong Liu, Hang Xu, Thomas Coudert, Pierangelo Renella, J Paul Finn, William Hsu, Kim-Lien Nguyen
arXiv AI
3d ago

From Compound Figures to Medical Multi-image Reasoning: Scaling Multimodal Large Language Models with Biomedical Literature

arXiv:2511.22232v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical inter...

By Zhen Chen, Yihang Fu, Rong Zhou, Serina Applebaum, Min Kyu Kim, Aidan Gilson, Morten Lee, Salahudeen Mirza, Gabriel Madera, Mauro Giuffre, Yuanting Pan, Roy Jiang, Hyunjae Kim, Hua Xu, Qingyu Chen