arXiv Machine Learning

Intersectional Disentangling of Temporal and Acquisition Bias in Fetal Ultrasound

arXiv:2605. 02942v2 Announce Type: replace Abstract: Fairness studies of medical imaging AI often explain subgroup performance gaps through under-representation in the training data.

arXiv Computer Vision
Sep 18

Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings

The paper introduces SonoCorpus, an open dataset of 456,963 ultrasound images with 1,626,085 expert masks from 53 public sources across 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on this data. SonoBase outperforms existing models (SAM2, MedSAM2, MedSAM3) on fifteen diverse evaluation datasets, matching specialist models and achieving clinically relevant accuracy for metrics such as ejection fraction, fetal head circumference, and gestational age. The authors provide full reproducibility resources, including checkpoints, optimizer states, and starter code, to enable community adoption and further development.

By Chao Qin, Fahad Shahbaz Khan, Salman Khan, Sarim Ather, Siddiq Anwar, Rao Muhammad Anwer, Shadab Khan
arXiv Machine Learning
Aug 27

FRAME: separating sampling variation from representational cause in medical imaging fairness

The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.

By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
arXiv AI
Jul 17

Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers

arXiv:2607. 14984v1 Announce Type: new Abstract: Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few samples that the resulting confidence intervals on per-subgroup performance are wider than the bias the audit is meant to detect.

By Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier
arXiv AI
Jul 9

CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation

arXiv:2603. 16551v2 Announce Type: replace-cross Abstract: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups.

By Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier
arXiv Computer Vision
Sep 4

Subgroup performance analysis of adaptation strategies for chest X-ray foundation models

The study examines how three parameter‑efficient adaptation methods—linear heads on the raw CLS token, an MLP, and an attention‑pooling module—affect pathology classification accuracy and subgroup fairness when applied to a frozen Rad‑DINO chest X‑ray encoder. Using the MIMIC‑CXR dataset, the authors evaluate eight pathologies across race, sex, and imaging‑view subgroups, finding that attention pooling yields the best overall performance and encodes protected attributes most strongly, yet higher performance does not consistently reduce subgroup disparities. The results show that attribute encoding strength and layer choice do not reliably predict fairness outcomes, indicating that fairness must be assessed directly for each task.

By Dhruv Gupta, Emma A. M. Stanley, Fabio De Sousa Ribeiro, Sujal R. Desai, Ben Glocker
arXiv AI
Jun 10

FADA: Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model

arXiv:2606. 11106v1 Announce Type: cross Abstract: A global shortage of trained sonographers limits prenatal ultrasound screening in low- and middle-income countries, where over half of pregnant women receive no skilled sonography.

By Mahmood Alzubaidi, Uzair Shah, Raden Muaz, Ines Abbes, Nader Mohammed, Abdullatif Magram, Khalid Alyafei, Mowafa Househ, Marco Agus
arXiv AI
Sep 16

Fairness at Every Intersection: Uncovering and Mitigating Intersectional Biases in Multimodal Clinical Predictions

The paper investigates intersectional biases in multimodal clinical predictions using Electronic Healthcare Records (EHR). It introduces datasets MIMIC-Eye1 and MIMIC-IV ED, applies unified text representations from pre‑trained clinical language models, and benchmarks bias mitigation at the intersectional subgroup level. Results show that subgroup‑specific mitigation is robust across datasets, subgroups, and embeddings, effectively addressing intersectional biases in multimodal settings.

By Ayaazuddin Mohammad, Kishore Sampath, Resmi Ramachandranpillai