The study evaluates a deep‑learning model for pediatric pneumonia detection across chest X‑ray datasets from three countries, assessing discrimination, calibration, operating‑point transport, shortcut signals, and limited‑label recovery. Using a frozen DenseNet121 ensemble trained on Guangzhou data, the model achieved high internal AUROC (0.976) but performance dropped when applied zero‑shot to Bangladesh (AUROC 0.798) and Vietnam (AUROC 0.742). Limited‑label adaptation with Platt recalibration restored sensitivity but introduced significant specificity variability, highlighting the need to evaluate multiple performance dimensions in cross‑dataset transport studies.
By Nazim-E-Alam
The study evaluates a deep‑learning model for pediatric pneumonia detection across three countries, testing not only discrimination but also probability calibration, fixed operating‑point transport, shortcut signals, and limited‑label recoverability. Using a DenseNet121 ensemble trained on Guangzhou data, the model achieved high internal AUROC (0.976) but performance dropped to 0.798 and 0.742 on Bangladeshi and Vietnamese datasets, respectively. Limited‑label adaptation with Platt recalibration restored sensitivity but introduced significant specificity variability, highlighting the need for comprehensive cross‑dataset evaluation.
Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely on pulmonary image content. This study evaluates...
arXiv:2606. 15910v2 Announce Type: replace Abstract: A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language priors.
By Reza Khanmohammadi, Kundan Thind, Mohammad M. Ghassemi
arXiv:2608. 15004v1 Announce Type: cross Abstract: Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment.
By Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta
arXiv:2606. 13211v1 Announce Type: new Abstract: AI systems are being deployed across medical imaging faster than their failure modes are understood.
By Omar Alshahrani, Muzammil Behzad
arXiv:2608.30467v1 Announce Type: new
Abstract: Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely...
By Abdullah Al Mamun, Md. Nasif Osman Khansur, Md Ashraful Hossen Akash, Md. Kishor Morol, Tze Hui Liew
arXiv:2608. 07857v1 Announce Type: cross Abstract: Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret.
By Fakrul Islam Tushar, Stephen Adamo, Geoffrey D. Rubin
arXiv:2608. 14766v1 Announce Type: cross Abstract: Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity.
By Simon Baur, Arne Schernich, Ekin B\"oke, Wojciech Samek, Jackie Ma
arXiv:2607. 25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases.
By Mateusz Koz{\l}owski
arXiv:2608.24881v1 Announce Type: cross
Abstract: Generative models are commonly ranked by Fr\'echet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary c...
By Hao Chen
arXiv:2607. 16888v1 Announce Type: cross Abstract: Deploying a medical imaging model that must later accommodate a modality it has never seen is a recurring practical problem: retraining the shared representation is expensive and destroys performance on the modalities already in service.
By Ranat Das Prangon, Istiaque Ahmed, Shajid Hasan Naim, Waseem Mustak Zisan, Hossain Md Shakhawat