In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroups. We study this pre-deployment fairness problem as an audit question: after a long-tailed multi-label CXR model is converted from scores into decisions, who is missed?
arXiv:2607. 28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups.
By Sparsh Roy, Samuel Girmachew, Nishita Chavan
The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.
By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
arXiv:2608. 14683v1 Announce Type: new Abstract: Given a patient's clinical findings, a diagnostic system ranks possible diseases and must decide when to endorse its first prediction or defer it for review.
By Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu
Subgroup performance differences are the standard evidence for fairness bias in medical imaging, and the usual response removes the demographic information that a model encodes. Here we introduce Fair...
Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.
By Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha