arXiv:2609.17545v1 Announce Type: new
Abstract: Deep learning models for cervical cytology are almost always evaluated as if every prediction must be acted upon, yet a screening system deployed along...
By Nisreen Albzour, Sarah S. Lam
The paper introduces a dual cross‑attention deep learning framework for automated grading of Cervical Intraepithelial Neoplasia (CIN) and prediction of Swede scores, using a newly released BUET Multi‑Center Colposcopy Dataset. The architecture fuses paired multimodal cervigrams and employs a custom composite loss to handle class imbalance, achieving 71.85% accuracy and 86.23% AUC‑ROC for three‑class CIN grading, and AUC‑ROC values between 75.7% and 88.4% for individual Swede score components. The total predicted Swede Score has a mean absolute error of 1.489, indicating potential for AI‑assisted colposcopy screening in resource‑limited settings.
By Dania Khan, Nuzhat Aisha Shaikh, Asfina Hassan Juicy, Raiyun Kabir, S M Shahida, Taufiq Hasan
arXiv:2607. 12075v1 Announce Type: cross Abstract: Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabilities, uncertainty estimation, and selective referral, particularly under dataset shift.
By Md. Sadibul Hasan Sadib, Md. Mohayminul Mukit, Rahmatul Kabir Rasel Sarker, Tahmid Alam Tamim, Md. Monir Hossain Shimul
arXiv:2608.22059v1 Announce Type: cross
Abstract: Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited....
By Xingtao Lin, Hangqi Ren, Caiwan Sun, You Chen
The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.
By Riya Deepak Shet, Chenxi Liang, Le Zhang
arXiv:2608. 14768v1 Announce Type: cross Abstract: Skin lesion classifiers can be confidently wrong on the cases that matter most, so knowing when a prediction should not be trusted is clinically as useful as the prediction.
By Leon Koole, Jiapan Guo, Matias Valdenegro-Toro
arXiv:2608. 08920v1 Announce Type: new Abstract: Perioperative risk prediction models are often limited by narrow surgical populations, incomplete intraoperative data, poor calibration, and limited interpretability.
By Shikhar Shukla, Cristina Barboi
arXiv:2608. 14866v1 Announce Type: cross Abstract: Objective: Small-sample molecular classification requires feature selectors that identify predictive, stable, and nonredundant subsets for binary and multiclass outcomes.
By Zardad Khan, Amjad Ali, Naz Gul, Sheema Gul, Saeed Aldahmani
The paper introduces a framework that distinguishes two causes of saturation in clinical prediction: a learner gap, where the model fails to use available information, and a measurement‑channel ceiling, where the recorded variables limit performance. It provides theoretical characterizations, finite‑sample diagnostics, and empirical audits across three large cohorts, showing that well‑tuned models approach the frontier while deficient learners leave large gaps. A PRISMA‑guided synthesis across 104 tasks reveals consistent channel‑level patterns, suggesting that improving the learner or the measurement channel can audit and potentially lift performance.
By Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan
arXiv:2608. 28063v1 Announce Type: new Abstract: Multi-organ ultrasound classifiers increasingly combine attention, mixture-of-experts routing, uncertainty gating, and evidential deep learning (EDL) objectives to address heterogeneous anatomy and acquisition.
By Yang Song, Pengbo Sun, Shichang Feng, Ye Zhu, Xin Xu, Ziran Wang
arXiv:2607. 16802v1 Announce Type: new Abstract: Deep survival models are evaluated almost exclusively by the concordance index (C-index), yet they are commonly trained using likelihood objectives such as the Cox partial likelihood, discrete-time negative log-likelihood, and DeepHit likelihood.
By Meixu Chen, Kai Wang, Jing Wang
arXiv:2607. 16317v1 Announce Type: cross Abstract: Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic.
By Medhansh Sharma