Improving Fairness of Large Language Model-Based ICU Mortality Prediction via Case-Based Prompting
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper examines how fairness conclusions in ICU mortality prediction using MIMIC-IV depend on the choice of metrics and the granularity of demographic analysis. It compares predictive-utility and subgroup-error metrics across various fairness interventions and introduces a lightweight adaptation strategy that balances ethnicity, gender, and insurance representation without conditioning on mortality outcomes. The study finds that different interventions can be evaluated differently across accuracy, sensitivity, and false-positive rate, and that marginal demographic summaries may hide heterogeneous error patterns within intersectional subgroups.
The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.
The paper evaluates machine learning models that predict retention and premature discontinuation in medication for opioid use disorder (MOUD). Using the Treatment Episode Data Set-Discharges (TEDS‑D) from 2015‑2019, the authors trained four models and examined overall performance as well as subgroup error rates by race, ethnicity, age, and sex. They also tested bias‑mitigation techniques, finding that these can reduce but not eliminate performance gaps without harming predictive accuracy.
The paper investigates intersectional biases in multimodal clinical predictions using Electronic Healthcare Records (EHR). It introduces datasets MIMIC-Eye1 and MIMIC-IV ED, applies unified text representations from pre‑trained clinical language models, and benchmarks bias mitigation at the intersectional subgroup level. Results show that subgroup‑specific mitigation is robust across datasets, subgroups, and embeddings, effectively addressing intersectional biases in multimodal settings.
FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.
arXiv:2608. 15254v1 Announce Type: new Abstract: Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI).