arXiv Machine Learning By Linus Juni, Aasa Feragen, Aditya Parikh

False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation

Read the original on arXiv Machine Learning →

arXiv:2607. 07852v1 Announce Type: cross Abstract: Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

FRAME: separating sampling variation from representational cause in medical imaging fairness

The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.

By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
arXiv AI
Sep 2

Causal Evidentiary Governance for High-Risk Machine Learning Systems

The paper proposes Causal Evidentiary Governance (CEG), a framework that requires regulated institutions to maintain a versioned directed acyclic graph (DAG) separating allowable from disallowed causal pathways in high‑risk machine learning systems. CEG introduces the Causal Harm Rate to quantify prediction variation due to disallowed pathways and pairs each decision with a signed Decision‑Evidence Packet (DEP) that cryptographically links the prediction to the DAG and path‑specific attributions, enabling efficient inclusion proofs via a Merkle tree. Empirical validation on synthetic credit data and the German Credit dataset demonstrates that CEG more clearly isolates causal effects than traditional fairness metrics and that a proof‑of‑concept implementation shows operational feasibility with manageable performance tradeoffs.

By Samah Kareem, Bar{\i}\c{s} \c{C}elikta\c{s}
arXiv AI
Jul 17

Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers

arXiv:2607. 14984v1 Announce Type: new Abstract: Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few samples that the resulting confidence intervals on per-subgroup performance are wider than the bias the audit is meant to detect.

By Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier