arXiv Machine Learning

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

The paper evaluates machine learning models that predict retention and premature discontinuation in medication for opioid use disorder (MOUD). Using the Treatment Episode Data Set-Discharges (TEDS‑D) from 2015‑2019, the authors trained four models and examined overall performance as well as subgroup error rates by race, ethnicity, age, and sex. They also tested bias‑mitigation techniques, finding that these can reduce but not eliminate performance gaps without harming predictive accuracy.

Hugging Face Trending Papers
Jul 9

FairSelect: A Systematic Evaluation of Multi-Level and Intersectional Algorithmic Fairness

Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for selecting fairness strategies, where disparities may arise across intersectional subgroups and across multiple stages of the modeling lifecycle.

arXiv Machine Learning
1d ago

Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction

The paper examines how fairness conclusions in ICU mortality prediction using MIMIC-IV depend on the choice of metrics and the granularity of demographic analysis. It compares predictive-utility and subgroup-error metrics across various fairness interventions and introduces a lightweight adaptation strategy that balances ethnicity, gender, and insurance representation without conditioning on mortality outcomes. The study finds that different interventions can be evaluated differently across accuracy, sensitivity, and false-positive rate, and that marginal demographic summaries may hide heterogeneous error patterns within intersectional subgroups.

By Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem, Syed Muhammad Ibne Zulfiker, Rishov Paul, Tanzima HAshem
arXiv AI
Aug 11

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

arXiv:2505. 16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability.

By Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
arXiv AI
Aug 20

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.

By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
arXiv AI
Aug 7

From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction

arXiv:2608. 05203v1 Announce Type: new Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning.

By Esra Zihni, Katryna Cisek, Hamzah Ziadeh, Hendrik Knoche, Robert Mikulik, John D. Kelleher