arXiv Machine Learning

Improving Fairness of Large Language Model-Based ICU Mortality Prediction via Case-Based Prompting

arXiv Machine Learning
1d ago

Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction

The paper examines how fairness conclusions in ICU mortality prediction using MIMIC-IV depend on the choice of metrics and the granularity of demographic analysis. It compares predictive-utility and subgroup-error metrics across various fairness interventions and introduces a lightweight adaptation strategy that balances ethnicity, gender, and insurance representation without conditioning on mortality outcomes. The study finds that different interventions can be evaluated differently across accuracy, sensitivity, and false-positive rate, and that marginal demographic summaries may hide heterogeneous error patterns within intersectional subgroups.

By Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem, Syed Muhammad Ibne Zulfiker, Rishov Paul, Tanzima HAshem
arXiv AI
2d ago

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.

By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
arXiv Machine Learning
Sep 22

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

The paper evaluates machine learning models that predict retention and premature discontinuation in medication for opioid use disorder (MOUD). Using the Treatment Episode Data Set-Discharges (TEDS‑D) from 2015‑2019, the authors trained four models and examined overall performance as well as subgroup error rates by race, ethnicity, age, and sex. They also tested bias‑mitigation techniques, finding that these can reduce but not eliminate performance gaps without harming predictive accuracy.

By Tongnian Wang, Carolina Vivas-Valencia, Cici Bauer, Yanmin Gong, Kim-Kwang Raymond Choo, Yuanxiong Guo
arXiv AI
Sep 16

Fairness at Every Intersection: Uncovering and Mitigating Intersectional Biases in Multimodal Clinical Predictions

The paper investigates intersectional biases in multimodal clinical predictions using Electronic Healthcare Records (EHR). It introduces datasets MIMIC-Eye1 and MIMIC-IV ED, applies unified text representations from pre‑trained clinical language models, and benchmarks bias mitigation at the intersectional subgroup level. Results show that subgroup‑specific mitigation is robust across datasets, subgroups, and embeddings, effectively addressing intersectional biases in multimodal settings.

By Ayaazuddin Mohammad, Kishore Sampath, Resmi Ramachandranpillai
arXiv AI
Aug 20

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.

By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
arXiv AI
Aug 11

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

arXiv:2505. 16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability.

By Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
arXiv Computation and Language
3d ago

Large Language Models are Approximate Survival Estimators

arXiv:2609.38181v1 Announce Type: new Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic...

By Juan M Zambrano Chaves, Peniel Argaw, Risa Ueno, Carlo Bifulco, Kristina Young, Rom Leidner, Tristan Naumann, Hoifung Poon
arXiv AI
Aug 28

Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset

The study evaluates a standalone large language model (LLM) versus a four‑step agentic pipeline for generating explanations of ICU mortality predictions on the eICU Demo dataset. XGBoost achieved an AUROC of 0.855 and an AUPRC of 0.332. In a 38‑case explanation subset, the standalone LLM produced one explanation with outcome leakage, while the agentic pipeline produced none; among 14 overlapping SHAP cases, the standalone LLM had higher SHAP alignment and direction consistency, whereas the agentic pipeline showed better guideline grounding, value specificity, and plausibility.

By Di Zhu, Chen Xie, Haoyun Zhang, Zihan Wei, Ziwei Wang, Jiazhao Shi, Ziyu Wang, Qiyang Xie