Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for selecting fairness strategies, where disparities may arise across intersectional subgroups and across multiple stages of the modeling lifecycle.
arXiv:2607. 08953v1 Announce Type: new Abstract: Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes.
By Nick Souligne, Isabella Mixton-Garcia, Vignesh Subbian
The paper examines how fairness conclusions in ICU mortality prediction using MIMIC-IV depend on the choice of metrics and the granularity of demographic analysis. It compares predictive-utility and subgroup-error metrics across various fairness interventions and introduces a lightweight adaptation strategy that balances ethnicity, gender, and insurance representation without conditioning on mortality outcomes. The study finds that different interventions can be evaluated differently across accuracy, sensitivity, and false-positive rate, and that marginal demographic summaries may hide heterogeneous error patterns within intersectional subgroups.
By Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem, Syed Muhammad Ibne Zulfiker, Rishov Paul, Tanzima HAshem
The paper investigates intersectional biases in multimodal clinical predictions using Electronic Healthcare Records (EHR). It introduces datasets MIMIC-Eye1 and MIMIC-IV ED, applies unified text representations from pre‑trained clinical language models, and benchmarks bias mitigation at the intersectional subgroup level. Results show that subgroup‑specific mitigation is robust across datasets, subgroups, and embeddings, effectively addressing intersectional biases in multimodal settings.
By Ayaazuddin Mohammad, Kishore Sampath, Resmi Ramachandranpillai
FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.
By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.
By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
The paper introduces the Explanation Consistency Score (ECS), a fairness‑aware metric that uses Jensen‑Shannon divergence to measure how similar attribution maps are across demographic subgroups. Applied to diabetic retinopathy screening, ECS is evaluated both overall and within disease severity levels. Results show that although predictive performance varies among ethnic groups, explanation consistency remains high and is not significantly linked to performance disparities, indicating that predictive fairness and explanation consistency assess different aspects of model behavior.
By Kerol Djoumessi, Philipp Berens
arXiv:2607. 16253v1 Announce Type: cross Abstract: Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world applications due to deficient external testing and fairness assessment.
By Rajveer Singh Pall, Sameer Yadav, Siddharth Bhalerao, Sourabh Sahu, Ritu Ahluwalia, Bhaskar Awadhiya
Subgroup performance differences are the standard evidence for fairness bias in medical imaging, and the usual response removes the demographic information that a model encodes. Here we introduce Fair...
arXiv:2512.19735v4 Announce Type: replace
Abstract: Accurately predicting mortality risk in intensive care unit (ICU) patients is critical for clinical decision-making. Large language models (LLMs) a...
By Gangxiong Zhang, Yongchao Long, Yuxi Zhou, Yong Zhang, Shenda Hong
arXiv:2609.07959v1 Announce Type: cross
Abstract: Machine learning models are widely used in clinical applications, social media, law enforcement and critical infrastructure. Verifying whether their...
By Francesca Panero, Ernst C. Wit, Marco Scutari
The paper examines how multi‑level fairness techniques—combining several bias‑mitigation steps—can reduce biases across patient demographics in health informatics. It reviews current literature, identifies gaps in implementation and reporting of health equity outcomes, and evaluates the role of reporting standards such as MINIMAR and TRIPOD in enhancing transparency. The authors conclude with recommendations to improve reporting transparency, broaden adoption of multi‑level fairness methods, and explicitly prioritize health equity in future research.
By Nick Souligne, Vignesh Subbian