Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for selecting fairness strategies, where disparities may arise across intersectional subgroups and across multiple stages of the modeling lifecycle.
arXiv:2604. 16450v2 Announce Type: replace-cross Abstract: Intersectional biases in healthcare data can produce compound disparities in clinical machine learning models, yet most fairness evaluations assess demographic attributes independently.
By Nick Souligne, Vignesh Subbian
The paper examines how fairness conclusions in ICU mortality prediction using MIMIC-IV depend on the choice of metrics and the granularity of demographic analysis. It compares predictive-utility and subgroup-error metrics across various fairness interventions and introduces a lightweight adaptation strategy that balances ethnicity, gender, and insurance representation without conditioning on mortality outcomes. The study finds that different interventions can be evaluated differently across accuracy, sensitivity, and false-positive rate, and that marginal demographic summaries may hide heterogeneous error patterns within intersectional subgroups.
By Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem, Syed Muhammad Ibne Zulfiker, Rishov Paul, Tanzima HAshem
arXiv:2606. 00656v1 Announce Type: cross Abstract: Ensuring fair and equitable treatment across diverse groups, particularly in multi-class classification tasks, poses a significant challenge due to the persistent biases inherent in machine learning models.
By Li Zhang, Yuyuan Li, XiaoHua Feng, Jiaming Zhang, Fengyuan Yu, Chaochao Chen
arXiv:2607. 16253v1 Announce Type: cross Abstract: Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world applications due to deficient external testing and fairness assessment.
By Rajveer Singh Pall, Sameer Yadav, Siddharth Bhalerao, Sourabh Sahu, Ritu Ahluwalia, Bhaskar Awadhiya
arXiv:2606. 20461v1 Announce Type: new Abstract: Machine learning models have been shown to exhibit discriminatory outcomes or degraded performance for individuals at the intersection of multiple sensitive attributes, such as race and gender.
By Bruno Scarone, Alfredo Viola, Ren\'ee J. Miller
arXiv:2602. 16794v2 Announce Type: replace-cross Abstract: Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored.
By Pengqi Liu, Zijun Yu, Mouloud Belbahri, Arthur Charpentier, Masoud Asgharian, Jesse C. Cresswell
The paper evaluates machine learning models that predict retention and premature discontinuation in medication for opioid use disorder (MOUD). Using the Treatment Episode Data Set-Discharges (TEDS‑D) from 2015‑2019, the authors trained four models and examined overall performance as well as subgroup error rates by race, ethnicity, age, and sex. They also tested bias‑mitigation techniques, finding that these can reduce but not eliminate performance gaps without harming predictive accuracy.
By Tongnian Wang, Carolina Vivas-Valencia, Cici Bauer, Yanmin Gong, Kim-Kwang Raymond Choo, Yuanxiong Guo
FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.
By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
The paper investigates intersectional biases in multimodal clinical predictions using Electronic Healthcare Records (EHR). It introduces datasets MIMIC-Eye1 and MIMIC-IV ED, applies unified text representations from pre‑trained clinical language models, and benchmarks bias mitigation at the intersectional subgroup level. Results show that subgroup‑specific mitigation is robust across datasets, subgroups, and embeddings, effectively addressing intersectional biases in multimodal settings.
By Ayaazuddin Mohammad, Kishore Sampath, Resmi Ramachandranpillai
arXiv:2609.07959v1 Announce Type: cross
Abstract: Machine learning models are widely used in clinical applications, social media, law enforcement and critical infrastructure. Verifying whether their...
By Francesca Panero, Ernst C. Wit, Marco Scutari
The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.
By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh