FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.
By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
arXiv:2607. 07717v1 Announce Type: new Abstract: In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroups.
By Ha-Hieu Pham, Hai-Dang Nguyen, Dang P. M. Cao, Thanh-Huy Nguyen, Min Xu, Trung-Nghia Le, Ulas Bagci, Huy-Hieu Pham
In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroups. We study this pre-deployment fairness problem as an audit question: after a long-tailed multi-label CXR model is converted from scores into decisions, who is missed?
arXiv:2607. 08953v1 Announce Type: new Abstract: Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes.
By Nick Souligne, Isabella Mixton-Garcia, Vignesh Subbian
The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.
By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
arXiv:2607. 28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups.
By Sparsh Roy, Samuel Girmachew, Nishita Chavan
arXiv:2605. 02942v2 Announce Type: replace Abstract: Fairness studies of medical imaging AI often explain subgroup performance gaps through under-representation in the training data.
By Aya Elgebaly, Joris Fournel, Benjamin Laine J{\o}nch Jurgensen, Kamil Mikolaj, Anders Christensen, Martin Tolsgaard, Claes Ladefoged, Aasa Feragen
arXiv:2605. 12895v2 Announce Type: replace-cross Abstract: Clinical decision-support systems are expert systems whose recommendations clinicians act on directly, yet they are usually cleared on one aggregate accuracy number from a held-out test set.
By Rohith Reddy Bellibatlu, Manpreet Singh, Yash Jajoo, Shyamal Lakhanpal, Abhishek Israni
Posted by Mike Schaekermann, Research Scientist, Google Research, and Ivor Horn, Chief Health Equity Officer & Director, Google Core Health equity is a major societal concern worldwide with disparities having many causes. These sources include limitations in access to healthcare, differences in clinical treatment, and even fundamental differences in the diagnostic technology.
By Google AI
arXiv:2604. 16450v2 Announce Type: replace-cross Abstract: Intersectional biases in healthcare data can produce compound disparities in clinical machine learning models, yet most fairness evaluations assess demographic attributes independently.
By Nick Souligne, Vignesh Subbian
arXiv:2603. 25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity).
By Jon-Paul Cacioli
arXiv:2606. 03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluation conditions has not been quantitatively characterized.
By Sangwon Baek, Kyu Yeon Hur, Kyunga Kim