arXiv Computer Vision

A multicenter benchmark and clinically structured metric for coronary CTA report generation

arXiv Machine Learning
Jul 30

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.

By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
arXiv AI
Jul 8

Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context

arXiv:2607. 05880v1 Announce Type: cross Abstract: Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone.

By Suneeta Mall, Vladimir Nekrasov, Ashnil Kumar, Sajith Karunasena, Aiden Nibali, Alix Bird, Mateo Diaz Shine, Jarrel Seah
arXiv Machine Learning
Aug 4

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

arXiv:2608. 00147v1 Announce Type: cross Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility.

By Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf, Lena Schmitzer, Lea Schumann, Jannik Kahmann, Friedrich Puttkammer, Johannes Moll, Jannik L\"ubberstedt, Zeineb Ben Chaaben, Anirudh Narayanan, Cosmin I. Bercea, Sebastian Ziegelmayer, Marcus R. Makowski, Daniel Rueckert, Lisa C. Adams, Keno K. Bressem
arXiv AI
Aug 5

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.

By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
arXiv AI
Aug 24

Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes

A multicenter study developed an AI-enabled electrocardiography (AI-ECG) model that predicts vessel-specific hemodynamically significant stenosis using coronary computed tomographic angiography (CCTA) as the reference. The model demonstrated strong discrimination in internal and external cohorts, including normal ECGs, and produced low-, intermediate-, and high-risk strata that correlated with stenosis severity and major adverse cardiovascular events. Calibration, decision curve analyses, and integration with guideline-based pre-test probability showed clinical utility, while waveform and attribution analyses revealed physiologically meaningful ECG features linked to high-risk predictions.

By Yujie Xiao, Qinghao Zhao, Gongzheng Tang, Hao Zhang, Zhuoran Kan, Deyun Zhang, Jun Li, Guangkun Nie, Xiaocheng Fang, Haoyu Wang, Shun Huang, Tong Liu, Jian Liu, Kangyin Chen, Shenda Hong
arXiv Computer Vision
Aug 27

Auditable CT Phenotyping Through Report-derived Radiological Observations

The study introduces Auditable CT Phenotyping (ACT), a method that uses report-derived radiological observations to predict 221 electronic-health-record phenotypes from CT scans. Trained on 38,317 patients and evaluated on 25,183 held‑out patients, ACT outperformed five vision‑language baselines and CT‑CLIP in both zero‑shot and linear probing settings. Analysis of the model’s probes revealed that a small set of observations, such as aortic and coronary calcification, dominated top predictions, but restricting the observation bank to clinician‑specified evidence improved interpretability without sacrificing accuracy.

By Riga Wu, Walter Witschey, Yicheng Li, Felix Barajas Ordonez, Keno K. Bressem, Lisa C. Adams, Gary E. Weissman, Li Shen, Christos Davatzikos, Eduardo Barbosa, Daniel Truhn, Tianyu Han
arXiv Computer Vision
Aug 26

Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

The study investigates how refined annotations versus model training changes affect pulmonary embolism segmentation performance. By re‑annotating 149 CT pulmonary angiography cases and evaluating two pretrained nnU-Net models, the authors find that improving annotation quality increases the Dice similarity coefficient (DSC) by 0.143–0.188, far exceeding the 0.028 DSC change from altering training datasets. A new human‑referenced benchmark model (nnPE) was trained and publicly released, though it performed below all annotators in paired comparisons.

By Qihang Sun, Zhongxiao Liu, Bailiang Jian, Shenman Qiu, Jingyuan Wang, Lei Zhang, Lixiang Xie, Jiazhen Pan, Christian Wachinger