The paper introduces SonoCorpus, an open dataset of 456,963 ultrasound images with 1,626,085 expert masks from 53 public sources across 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on this data. SonoBase outperforms existing models (SAM2, MedSAM2, MedSAM3) on fifteen diverse evaluation datasets, matching specialist models and achieving clinically relevant accuracy for metrics such as ejection fraction, fetal head circumference, and gestational age. The authors provide full reproducibility resources, including checkpoints, optimizer states, and starter code, to enable community adoption and further development.
By Chao Qin, Fahad Shahbaz Khan, Salman Khan, Sarim Ather, Siddiq Anwar, Rao Muhammad Anwer, Shadab Khan
The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.
By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
Subgroup performance differences are the standard evidence for fairness bias in medical imaging, and the usual response removes the demographic information that a model encodes. Here we introduce Fair...
arXiv:2607. 14984v1 Announce Type: new Abstract: Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few samples that the resulting confidence intervals on per-subgroup performance are wider than the bias the audit is meant to detect.
By Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier
arXiv:2608. 28063v1 Announce Type: new Abstract: Multi-organ ultrasound classifiers increasingly combine attention, mixture-of-experts routing, uncertainty gating, and evidential deep learning (EDL) objectives to address heterogeneous anatomy and acquisition.
By Yang Song, Pengbo Sun, Shichang Feng, Ye Zhu, Xin Xu, Ziran Wang
arXiv:2603. 16551v2 Announce Type: replace-cross Abstract: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups.
By Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier
Self-supervised learning (SSL) can reduce the need for labelled medical images, but the choice of pretext objective remains unclear for lung ultrasound (LUS). Contrastive learning, masked reconstructi...
The study examines how three parameter‑efficient adaptation methods—linear heads on the raw CLS token, an MLP, and an attention‑pooling module—affect pathology classification accuracy and subgroup fairness when applied to a frozen Rad‑DINO chest X‑ray encoder. Using the MIMIC‑CXR dataset, the authors evaluate eight pathologies across race, sex, and imaging‑view subgroups, finding that attention pooling yields the best overall performance and encodes protected attributes most strongly, yet higher performance does not consistently reduce subgroup disparities. The results show that attribute encoding strength and layer choice do not reliably predict fairness outcomes, indicating that fairness must be assessed directly for each task.
By Dhruv Gupta, Emma A. M. Stanley, Fabio De Sousa Ribeiro, Sujal R. Desai, Ben Glocker
arXiv:2509. 19671v3 Announce Type: replace Abstract: Public datasets of Chest X-Rays (CXRs) have long been a popular benchmark for developing machine learning (ML) computer vision models in healthcare.
By Andrew Wang, Jiashuo Zhang, Michael Oberst
arXiv:2606. 11106v1 Announce Type: cross Abstract: A global shortage of trained sonographers limits prenatal ultrasound screening in low- and middle-income countries, where over half of pregnant women receive no skilled sonography.
By Mahmood Alzubaidi, Uzair Shah, Raden Muaz, Ines Abbes, Nader Mohammed, Abdullatif Magram, Khalid Alyafei, Mowafa Househ, Marco Agus
The paper investigates intersectional biases in multimodal clinical predictions using Electronic Healthcare Records (EHR). It introduces datasets MIMIC-Eye1 and MIMIC-IV ED, applies unified text representations from pre‑trained clinical language models, and benchmarks bias mitigation at the intersectional subgroup level. Results show that subgroup‑specific mitigation is robust across datasets, subgroups, and embeddings, effectively addressing intersectional biases in multimodal settings.
By Ayaazuddin Mohammad, Kishore Sampath, Resmi Ramachandranpillai
arXiv:2607. 07852v1 Announce Type: cross Abstract: Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy.
By Linus Juni, Aasa Feragen, Aditya Parikh