arXiv Machine Learning

When Average Calibration Fails: Site-Conditional Federated Conformal Risk Control

arXiv:2606. 20115v3 Announce Type: replace Abstract: Conformal risk control (CRC) provides distribution-free segmentation guarantees by calibrating a prediction-set threshold on held-out data.

arXiv Machine Learning
5d ago

Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift

The paper introduces a missingness‑aware conformal calibration method for mortality prediction that accounts for cross‑hospital distribution shifts. By selecting a measurement on an independent sample, grouping patients by whether that measurement is recorded, and applying Mondrian calibration within each group, the method avoids reusing calibration outcomes. Experiments on eICU and MIMIC‑IV data show that, compared to pooled calibration, it reduces the worst‑group coverage gap by a median of 1.9 percentage points across six settings, though the benefit varies with predictor and hospital.

By Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai
arXiv Machine Learning
4d ago

ReCIRC: Rectified Conformal Risk Control

arXiv:2609.38112v1 Announce Type: cross Abstract: Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or misse...

By Bruno Marcondes e Resende, Helton Graziadei, Thiago Rodrigo Ramos, Rafael Izbicki
arXiv AI
Aug 20

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

The paper presents a distribution‑free risk‑control framework that provides organ‑specific recall guarantees for frozen multi‑organ CT segmentation models. It calibrates per‑organ thresholds for an AMOS‑trained nnU‑Net, audits its transfer to RAOS, and estimates local re‑certification costs using case‑level voxel false‑negative rates. The study compares Risk‑Controlling Prediction Sets (RCPS) and Conformal Risk Control (CRC), noting that RCPS offers high‑probability control of population‑mean risk while CRC provides weaker expectation control, and evaluates the effectiveness of the Waudby–Smith–Ramdas betting bound versus Hoeffding–Bentkus bounds for re‑certification. "whyItMatters":"The work demonstrates how to maintain organ‑level recall guarantees when deploying segmentation models across different clinical domains, highlighting the trade‑offs between threshold conservatism and re‑certification effort."

By Souraj Adhikary, Negar Chabi, Andre Mastmeyer
Hugging Face Trending Papers
Aug 18

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

The paper introduces a distribution‑free risk control method that provides organ‑specific recall guarantees for frozen segmentation models. It calibrates per‑organ thresholds on an AMOS‑trained nnU‑Net, audits transfer to RAOS, and estimates local re‑certification cost using case‑level voxel false‑negative rates. The study compares Risk‑Controlling Prediction Sets (RCPS) and Conformal Risk Control (CRC), noting that RCPS offers high‑probability control of population‑mean risk while CRC provides weaker expectation control, and evaluates the effectiveness of the Waudby‑Smith‑Ramdas betting bound versus Hoeffding‑Bentkus bounds for re‑certification of Tier‑1 organs.

arXiv Machine Learning
Sep 10

Conditional Validity for Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group

The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.

By Melika Baghi
arXiv Machine Learning
Jul 28

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

arXiv:2607. 22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning.

By Pranav Kaliaperumal, Manisha Kaliaperumal