The study evaluates a deep‑learning model for pediatric pneumonia detection across chest X‑ray datasets from three countries, assessing discrimination, calibration, operating‑point transport, shortcut signals, and limited‑label recovery. Using a frozen DenseNet121 ensemble trained on Guangzhou data, the model achieved high internal AUROC (0.976) but performance dropped when applied zero‑shot to Bangladesh (AUROC 0.798) and Vietnam (AUROC 0.742). Limited‑label adaptation with Platt recalibration restored sensitivity but introduced significant specificity variability, highlighting the need to evaluate multiple performance dimensions in cross‑dataset transport studies.
By Nazim-E-Alam
Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely on pulmonary image content. This study evaluates...
arXiv:2608.30467v1 Announce Type: new
Abstract: Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely...
By Abdullah Al Mamun, Md. Nasif Osman Khansur, Md Ashraful Hossen Akash, Md. Kishor Morol, Tze Hui Liew
arXiv:2607. 09305v1 Announce Type: cross Abstract: Chest radiography (CXR) remains the most widely used thoracic imaging modality, yet expert interpretation is constrained by a severe shortage of radiologists in Thailand and across Southeast Asia.
By Isarun Chamveha, Tretap Promwiset, Napat Wanchaitanawong, Trongtum Tongdee, Pairash Saiviroonporn, Warasinee Chaisangmongkon
Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.
By Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha
arXiv:2606. 12824v1 Announce Type: cross Abstract: AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends local acceptance testing and ongoing drift monitoring, and the ACR Assess-AI registry monitors AI outputs using DICOM metadata for context.
By Daniel Soliman
arXiv:2609.37848v1 Announce Type: cross
Abstract: Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even...
By Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala
arXiv:2608.22059v1 Announce Type: cross
Abstract: Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited....
By Xingtao Lin, Hangqi Ren, Caiwan Sun, You Chen
The study evaluates the robustness of medical vision‑language models for tuberculosis screening on chest X‑rays by testing them across multiple datasets, prompts, and evaluation settings. Three specialized models (BioMedCLIP, CheXficient, MedSigLIP) and a general OpenCLIP model were audited on 12,200 images, producing 244,000 model–image–prompt scores. Results show that no model consistently outperforms others across all cohorts and reliability criteria, with prompt changes and control group composition significantly affecting AUROC, and that high training‑set performance does not reliably transfer to external cohorts.
By Mushir Akhtar, M. Tanveer, Mohd. Arshad
arXiv:2608.29348v1 Announce Type: new
Abstract: Background: Patient details and acquisition metadata are important for clinical decisions, image quality control, and automated research pipelines, but...
By Jakob Wasserthal, Joshy Cyriac, Michael Bach, Kimia Mozahheb Yousefi, Minh-Son To, M\'at\'e Sik, C\'edric H\'emon, Thomas Weikert, Martin Segeroth
arXiv:2607. 07219v1 Announce Type: cross Abstract: Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous.
By Alejandro Vergara-Richart (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain, Universitat Polit\`ecnica de Val\`encia, Val\`encia, Spain), Xavier Rafael-Palou (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), Almudena Fuster-Matanzo (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), Ignacio Iborra Roncales (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), \'Angel Alberich-Bayarri (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), Ana Jim\'enez-Pastor (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain)
The study investigates how image‑domain Poisson perturbations affect the classification of ischemic core/penumbra in non‑contrast CT slices. Using a CPAISD cohort, the authors compared direct ResNet‑18 classification with a denoising‑then‑classifying pipeline and found that denoising generally reduced performance. A prospective experiment with joint denoising‑classification models showed no statistically significant advantage over direct noisy classification, indicating limited robustness to Poisson noise.
By Rhea Ghosal, Ronok Ghosal, Eileen Lou