arXiv Computer Vision By Nisreen Albzour, Sarah S. Lam

Reliability-Aware Hybrid-K Ensemble Selection for Cervical Cytology Classification: Integrating Discrimination, Calibration, and Selective Prediction

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

arXiv Computer Vision
5d ago

A Dual Cross-Attention Framework for Colposcopic CIN Grading and Swede Score Prediction Using a New Multi-Center Dataset

The paper introduces a dual cross‑attention deep learning framework for automated grading of Cervical Intraepithelial Neoplasia (CIN) and prediction of Swede scores, using a newly released BUET Multi‑Center Colposcopy Dataset. The architecture fuses paired multimodal cervigrams and employs a custom composite loss to handle class imbalance, achieving 71.85% accuracy and 86.23% AUC‑ROC for three‑class CIN grading, and AUC‑ROC values between 75.7% and 88.4% for individual Swede score components. The total predicted Swede Score has a mean absolute error of 1.489, indicating potential for AI‑assisted colposcopy screening in resource‑limited settings.

By Dania Khan, Nuzhat Aisha Shaikh, Asfina Hassan Juicy, Raiyun Kabir, S M Shahida, Taufiq Hasan
arXiv AI
Jul 15

Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation

arXiv:2607. 12075v1 Announce Type: cross Abstract: Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabilities, uncertainty estimation, and selective referral, particularly under dataset shift.

By Md. Sadibul Hasan Sadib, Md. Mohayminul Mukit, Rahmatul Kabir Rasel Sarker, Tahmid Alam Tamim, Md. Monir Hossain Shimul
arXiv Computer Vision
Aug 27

Reliability analysis for BraTS-GoAT segmentation: a controlled robustness study of deep-ensemble uncertainty

The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.

By Riya Deepak Shet, Chenxi Liang, Le Zhang