Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.
arXiv:2608. 14705v1 Announce Type: cross Abstract: Hyperparameter optimization (HPO) can materially affect the performance of deep learning (DL) image classifiers, but there is little empirical guidance on how to derive the validation signal that drives it, especially for the small sample sizes common in fields such as medical imaging.
Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely on pulmonary image content. This study evaluates...
arXiv:2608. 16709v1 Announce Type: cross Abstract: A radiologist reading a model's output faces two problems.
The paper introduces a label‑free method called AURCC for selecting the best foundational model for medical image classification when the target domain lacks labels. AURCC uses a pseudo‑label discrepancy computed by the SUDO framework to score models without fine‑tuning. Experiments on chest X‑ray data across three inter‑hospital shifts show that AURCC closely matches the true model ranking, outperforming simple source‑accuracy baselines especially when source data are limited.
arXiv:2608.30467v1 Announce Type: new Abstract: Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely...