arXiv Machine Learning By Sparsh Roy, Samuel Girmachew, Nishita Chavan

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

Read the original on arXiv Machine Learning →

arXiv:2607. 28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

The paper reports that counterfactual fairness audits of clinical language‑model agents are unreliable without accounting for a per‑action instability floor. By repeatedly running identical vignettes, the authors found that actions changed 8.7% of the time, with instability varying eightfold across actions. A second model confirmed a pooled floor of 6.7%, showing that any reported fairness estimate lacking this floor cannot be interpreted as evidence of disparity.

By Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar, Rahul Joshi