arXiv AI

Which Anatomy Matters Under Limited Labels? A Data-Efficient Anatomy-Aware Benchmark for Cardiac Pathology Prediction

arXiv:2606. 06509v1 Announce Type: cross Abstract: Numerous medical imaging problems must be solved under limited labels and constrained compute, yet it remains unclear whether performance gains are driven mainly by more expressive models or by better representation of clinically meaningful anatomy.

arXiv Computer Vision
Sep 14

SV-Cine: Diagnosis-Conditioned Segmentation of Single Ventricle Physiology via Generative Data Augmentation

The paper introduces SV-Cine, a cardiac MRI segmentation framework tailored for single ventricle physiology (SVP). It combines a generative data augmentation pipeline that creates synthetic 3D cardiac meshes and MRI, with a diagnosis-conditioned adaptation of the CineMA foundation model that uses patient-level diagnostic information to improve segmentation. Evaluations on an internal cohort show high Dice scores for left and right ventricles, outperforming nnU-Net, and demonstrate that incorporating diagnosis priors can adapt a pretrained model to specialized SVP tasks.

By Lila Cunge, Yuehong Liu, Hang Xu, Thomas Coudert, Pierangelo Renella, J Paul Finn, William Hsu, Kim-Lien Nguyen
Hugging Face Trending Papers
Sep 2

Learning from Scarce Labels: Multi-View Echocardiography for Ejection Fraction Prediction

The paper introduces the first publicly available dataset for predicting left ventricular ejection fraction (EF) from parasternal long-axis (PLAX) echocardiography, comprising over 25,000 labeled videos generated through a novel data‑generation strategy that correlates clinical notes with echocardiographic videos. Using this dataset, the authors train a reproducible PLAX‑EF model that achieves a mean absolute error (MAE) of 6.86%, comparable to the clinical standard of apical four‑chamber (A4C) methods. They further show that combining PLAX and A4C predictions via simple late fusion reduces MAE to 6.37%, highlighting the benefit of multi‑view integration, and they release the dataset, models, and demos for community use.

arXiv Machine Learning
Sep 4

Learning from Scarce Labels: Multi-View Echocardiography for Ejection Fraction Prediction

The paper introduces the first publicly available dataset of over 25,000 parasternal long‑axis (PLAX) echocardiography videos labeled for left ventricular ejection fraction (EF), created through a novel data‑generation strategy that links clinical notes to video data. Using this dataset, the authors train a reproducible PLAX‑based EF model that achieves a mean absolute error (MAE) of 6.86%, comparable to the clinical standard of apical four‑chamber (A4C) methods. They further show that simple late fusion of PLAX and A4C predictions reduces MAE to 6.37%, highlighting the benefit of multi‑view integration, and release the dataset, models, and demos publicly.

By Zhiyuan Gao, Dominic Yurk, Yaser S. Abu-Mostafa
arXiv AI
Jul 14

A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation

arXiv:2607. 11287v1 Announce Type: cross Abstract: Comprehensive quantification of cardiac structures from computed tomography (CT) remains limited not by data availability but by the scalability of measurements, which makes routine use impractical.

By Pooya Mohammadi Kazaj, Leo Fridolin Weber, Wen Xie, Seyed Amir Ahmad Safavi-Naini, Anselm Stark, Giovanni Baj, Ali Mokhtari, Toshiya Yoshida, Christoph Ryffel, Taishi Okuno, Yoshihiro Akashi, Ronny R. Buechel, Thomas Pilgrim, Waldo Valenzuela, George C. M. Siontis, Xiaowei Xu, Moritz Hundertmark, Stephan Windecker, Christoph Grani, Isaac Shiri
arXiv Machine Learning
Aug 18

CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration

arXiv:2608. 16268v1 Announce Type: cross Abstract: Medical foundation models improve generalization when training AI models with limited labeled data, but remain confined to a single specialty, such as pathology or radiology, and to either sparse or dense outputs, such as classification or segmentation.

By J. Raphael Sch\"afer, Kai Geissler, Till Nicke, Chiara Tappermann, Karoline Heber, Eike Petersen, Habib Mergan, Lars Ole Schwen, Nick Weiss, Annika Gerken, Jan Hendrik Moltz, Tom Bisson, Isil Dogan O, Tim-Rasmus Kiehl, Norman Zerbe, Sefer Elezkurtaj, Robin S. Mayer, Nadine Flinner, Peter Wild, Isabel Dahm, Felix Peisen, Heinrich von Busch, Robert Grimm, Sebastian Arndt, Lisa Siegler, Matthias Stefan May, Antje Prasse, Natalia Artysh, Fabian Kiessling, Johannes Lotz
arXiv Computer Vision
Sep 2

CMRVision: A Foundation Model for Cardiac MR Image Analysis

CMRVision is a cardiac magnetic resonance (CMR) foundation model trained with DINOv3-style self‑supervised learning on 36 million multi‑center, multi‑sequence CMR images. It outperforms prior natural‑image, medical‑image, supervised, and CMR baselines on multi‑task segmentation (cine, LGE, mapping) and cine view classification, achieving Dice scores of 0.940–0.967 for LV and 0.855–0.905 for myocardium, and a zero‑shot Dice of 0.692 on unseen LGE long‑axis views. The model demonstrates robust cross‑view generalization and highest average accuracy (0.906) for cine view classification.

By Athira J. Jacob, Puneet Sharma, Daniel Rueckert
arXiv Machine Learning
Aug 31

Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations

The paper introduces SPAR‑Bench, a set of eight probes designed to test whether medical vision models can reason about anatomy in abdominal CT scans. Experiments across five architectures and three foundation models—both frozen and fine‑tuned—show that while models can recall canonical organ locations, they fail to perform relational reasoning or spatial comparisons within a patient, even under zero‑shot transfer. The study also demonstrates that pooled probing underestimates a model’s relational capabilities and that open‑weight multimodal large language models perform poorly on these tasks.

By Naren Akash, Neeraja Ramanan
arXiv AI
Jul 23

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

arXiv:2607. 20274v1 Announce Type: cross Abstract: Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure.

By Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather, Daniel Truhn