arXiv Machine Learning

When Calibration Fails the Vulnerable Hospital: Federated Conformal Risk Control via Risk-Curve Shrinkage

arXiv:2606. 20115v1 Announce Type: new Abstract: Conformal risk control (CRC) provides distribution-free guarantees on segmentation quality by calibrating a prediction-set threshold on held-out data.

arXiv Machine Learning
5d ago

Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift

The paper introduces a missingness‑aware conformal calibration method for mortality prediction that accounts for cross‑hospital distribution shifts. By selecting a measurement on an independent sample, grouping patients by whether that measurement is recorded, and applying Mondrian calibration within each group, the method avoids reusing calibration outcomes. Experiments on eICU and MIMIC‑IV data show that, compared to pooled calibration, it reduces the worst‑group coverage gap by a median of 1.9 percentage points across six settings, though the benefit varies with predictor and hospital.

By Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai
arXiv Machine Learning
4d ago

ReCIRC: Rectified Conformal Risk Control

arXiv:2609.38112v1 Announce Type: cross Abstract: Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or misse...

By Bruno Marcondes e Resende, Helton Graziadei, Thiago Rodrigo Ramos, Rafael Izbicki
arXiv AI
Aug 20

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

The paper presents a distribution‑free risk‑control framework that provides organ‑specific recall guarantees for frozen multi‑organ CT segmentation models. It calibrates per‑organ thresholds for an AMOS‑trained nnU‑Net, audits its transfer to RAOS, and estimates local re‑certification costs using case‑level voxel false‑negative rates. The study compares Risk‑Controlling Prediction Sets (RCPS) and Conformal Risk Control (CRC), noting that RCPS offers high‑probability control of population‑mean risk while CRC provides weaker expectation control, and evaluates the effectiveness of the Waudby–Smith–Ramdas betting bound versus Hoeffding–Bentkus bounds for re‑certification. "whyItMatters":"The work demonstrates how to maintain organ‑level recall guarantees when deploying segmentation models across different clinical domains, highlighting the trade‑offs between threshold conservatism and re‑certification effort."

By Souraj Adhikary, Negar Chabi, Andre Mastmeyer
arXiv Machine Learning
Jul 30

CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

arXiv:2607. 26752v1 Announce Type: new Abstract: Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning.

By Behraj Khan, Shabir Ahmad, Syed Ahmad Chan Bukhari, Tahir Qasim Syed
Hugging Face Trending Papers
Aug 18

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

The paper introduces a distribution‑free risk control method that provides organ‑specific recall guarantees for frozen segmentation models. It calibrates per‑organ thresholds on an AMOS‑trained nnU‑Net, audits transfer to RAOS, and estimates local re‑certification cost using case‑level voxel false‑negative rates. The study compares Risk‑Controlling Prediction Sets (RCPS) and Conformal Risk Control (CRC), noting that RCPS offers high‑probability control of population‑mean risk while CRC provides weaker expectation control, and evaluates the effectiveness of the Waudby‑Smith‑Ramdas betting bound versus Hoeffding‑Bentkus bounds for re‑certification of Tier‑1 organs.

arXiv Machine Learning
Sep 21

Available Guardrails: Certifying Selective Prediction across ML Systems

The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.

By Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky
arXiv AI
Jul 15

Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI

arXiv:2505. 22108v4 Announce Type: replace-cross Abstract: Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is limited by privacy concerns, heterogeneous institutional compliance, and resource disparities; standard differential privacy (DP) applies uniform noise to all clients, penalizing well-compliant or under-resourced institutions.

By Santhosh Parampottupadam, Melih Co\c{s}\u{g}un, Sarthak Pati, Maximilian Zenk, Saikat Roy, Dimitrios Bounias, Benjamin Hamm, Sinem Sav, Ralf Floca, Klaus Maier-Hein