Subgroup performance differences are the standard evidence for fairness bias in medical imaging, and the usual response removes the demographic information that a model encodes. Here we introduce Fair...
arXiv:2607. 14984v1 Announce Type: new Abstract: Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few samples that the resulting confidence intervals on per-subgroup performance are wider than the bias the audit is meant to detect.
By Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier
The paper introduces the Explanation Consistency Score (ECS), a fairness‑aware metric that uses Jensen‑Shannon divergence to measure how similar attribution maps are across demographic subgroups. Applied to diabetic retinopathy screening, ECS is evaluated both overall and within disease severity levels. Results show that although predictive performance varies among ethnic groups, explanation consistency remains high and is not significantly linked to performance disparities, indicating that predictive fairness and explanation consistency assess different aspects of model behavior.
By Kerol Djoumessi, Philipp Berens
arXiv:2607. 07852v1 Announce Type: cross Abstract: Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy.
By Linus Juni, Aasa Feragen, Aditya Parikh
arXiv:2607. 08953v1 Announce Type: new Abstract: Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes.
By Nick Souligne, Isabella Mixton-Garcia, Vignesh Subbian
arXiv:2604. 16450v2 Announce Type: replace-cross Abstract: Intersectional biases in healthcare data can produce compound disparities in clinical machine learning models, yet most fairness evaluations assess demographic attributes independently.
By Nick Souligne, Vignesh Subbian