arXiv AI

Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection

arXiv Computer Vision
Aug 26

Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

The study investigates how refined annotations versus model training changes affect pulmonary embolism segmentation performance. By re‑annotating 149 CT pulmonary angiography cases and evaluating two pretrained nnU-Net models, the authors find that improving annotation quality increases the Dice similarity coefficient (DSC) by 0.143–0.188, far exceeding the 0.028 DSC change from altering training datasets. A new human‑referenced benchmark model (nnPE) was trained and publicly released, though it performed below all annotators in paired comparisons.

By Qihang Sun, Zhongxiao Liu, Bailiang Jian, Shenman Qiu, Jingyuan Wang, Lei Zhang, Lixiang Xie, Jiazhen Pan, Christian Wachinger
arXiv AI
Aug 24

Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes

A multicenter study developed an AI-enabled electrocardiography (AI-ECG) model that predicts vessel-specific hemodynamically significant stenosis using coronary computed tomographic angiography (CCTA) as the reference. The model demonstrated strong discrimination in internal and external cohorts, including normal ECGs, and produced low-, intermediate-, and high-risk strata that correlated with stenosis severity and major adverse cardiovascular events. Calibration, decision curve analyses, and integration with guideline-based pre-test probability showed clinical utility, while waveform and attribution analyses revealed physiologically meaningful ECG features linked to high-risk predictions.

By Yujie Xiao, Qinghao Zhao, Gongzheng Tang, Hao Zhang, Zhuoran Kan, Deyun Zhang, Jun Li, Guangkun Nie, Xiaocheng Fang, Haoyu Wang, Shun Huang, Tong Liu, Jian Liu, Kangyin Chen, Shenda Hong
arXiv Computer Vision
Aug 27

Auditable CT Phenotyping Through Report-derived Radiological Observations

The study introduces Auditable CT Phenotyping (ACT), a method that uses report-derived radiological observations to predict 221 electronic-health-record phenotypes from CT scans. Trained on 38,317 patients and evaluated on 25,183 held‑out patients, ACT outperformed five vision‑language baselines and CT‑CLIP in both zero‑shot and linear probing settings. Analysis of the model’s probes revealed that a small set of observations, such as aortic and coronary calcification, dominated top predictions, but restricting the observation bank to clinician‑specified evidence improved interpretability without sacrificing accuracy.

By Riga Wu, Walter Witschey, Yicheng Li, Felix Barajas Ordonez, Keno K. Bressem, Lisa C. Adams, Gary E. Weissman, Li Shen, Christos Davatzikos, Eduardo Barbosa, Daniel Truhn, Tianyu Han
arXiv Computer Vision
Sep 7

Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery

The study evaluates a deep‑learning model for pediatric pneumonia detection across chest X‑ray datasets from three countries, assessing discrimination, calibration, operating‑point transport, shortcut signals, and limited‑label recovery. Using a frozen DenseNet121 ensemble trained on Guangzhou data, the model achieved high internal AUROC (0.976) but performance dropped when applied zero‑shot to Bangladesh (AUROC 0.798) and Vietnam (AUROC 0.742). Limited‑label adaptation with Platt recalibration restored sensitivity but introduced significant specificity variability, highlighting the need to evaluate multiple performance dimensions in cross‑dataset transport studies.

By Nazim-E-Alam
arXiv AI
Jun 2

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

arXiv:2604. 15231v2 Announce Type: replace Abstract: Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT).

By M\'elanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Michael Krauthammer, Michael Moor
arXiv AI
Sep 25

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.

By Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha
arXiv Machine Learning
Jul 30

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.

By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
arXiv Machine Learning
Sep 22

ORION-CMR: On-scanner Reporting with Integrated Foundation Model for End-to-End Cardiac MRI Analysis and Interpretation

ORION‑CMR is a scanner‑native, end‑to‑end foundation model for cardiac MRI that performs sequence classification, ventricular function assessment, LGE detection, disease classification, and generates reports in about 90 seconds. Trained on 12.9 million images, it outperformed supervised baselines and a prior CMR foundation model, achieving state‑of‑the‑art LGE classification and scar segmentation. In a multi‑vendor clinical cohort, it reached an AUC of 0.96 for normal‑vs‑abnormal detection and 0.88 for multiclass disease classification, with generated reports agreeing 81.4% with expert interpretation.

By Omer Burak Demirel, Kelly K. Horst, Alessio Perazzolo, Elisa Bruno, Kenan Kaya, Rongzhen Ouyang, Enas Ahmed, Jouke Smink, Spencer L. Waddle, Zainudeen Kallumpurath, Tzu Cheng Chao, Dinghui Wang, Steve G. Langer, Timothy L. Kline, Panagiotis Korfiatis, Jacinta Browne, Ivana Isgum, Tim Leiner
arXiv Computer Vision
Sep 15

Mobile CT Services for Rural, Regional, and Remote Areas: Current Practice and Future Integration with Telehealth and Regulatory-Authorised AI

arXiv:2609.14347v1 Announce Type: new Abstract: Computed tomography (CT) plays an essential role in clinical workflow to improve patient outcomes. However, access to CT imaging and specialist interpr...

By Zhicheng Lu, Md Zahid Islam, M Mamun Huda, Kristie Sweeney, Shayne Chau, Oliver Mulcock, Corey Hemopo, Catherine Keniry, Mohammad Ali Moni