arXiv AI

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

The paper presents a distribution‑free risk‑control framework that provides organ‑specific recall guarantees for frozen multi‑organ CT segmentation models. It calibrates per‑organ thresholds for an AMOS‑trained nnU‑Net, audits its transfer to RAOS, and estimates local re‑certification costs using case‑level voxel false‑negative rates. The study compares Risk‑Controlling Prediction Sets (RCPS) and Conformal Risk Control (CRC), noting that RCPS offers high‑probability control of population‑mean risk while CRC provides weaker expectation control, and evaluates the effectiveness of the Waudby–Smith–Ramdas betting bound versus Hoeffding–Bentkus bounds for re‑certification. "whyItMatters":"The work demonstrates how to maintain organ‑level recall guarantees when deploying segmentation models across different clinical domains, highlighting the trade‑offs between threshold conservatism and re‑certification effort."

Hugging Face Trending Papers
Aug 18

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

The paper introduces a distribution‑free risk control method that provides organ‑specific recall guarantees for frozen segmentation models. It calibrates per‑organ thresholds on an AMOS‑trained nnU‑Net, audits transfer to RAOS, and estimates local re‑certification cost using case‑level voxel false‑negative rates. The study compares Risk‑Controlling Prediction Sets (RCPS) and Conformal Risk Control (CRC), noting that RCPS offers high‑probability control of population‑mean risk while CRC provides weaker expectation control, and evaluates the effectiveness of the Waudby‑Smith‑Ramdas betting bound versus Hoeffding‑Bentkus bounds for re‑certification of Tier‑1 organs.

arXiv Machine Learning
4d ago

ReCIRC: Rectified Conformal Risk Control

arXiv:2609.38112v1 Announce Type: cross Abstract: Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or misse...

By Bruno Marcondes e Resende, Helton Graziadei, Thiago Rodrigo Ramos, Rafael Izbicki
arXiv Machine Learning
Jul 28

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

arXiv:2607. 22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning.

By Pranav Kaliaperumal, Manisha Kaliaperumal
arXiv Machine Learning
Jun 9

A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

arXiv:2606. 08517v1 Announce Type: new Abstract: Selective predictors answer on confident inputs and abstain elsewhere; deploying one safely needs a single finite-sample certificate that simultaneously upper-bounds the selected risk, lower-bounds the acceptance probability $\pacc$ above a floor $\pmin$, and lower-bounds the deployment utility.

By Xiaoli Yu, Jiamiao Liu
arXiv Machine Learning
Jul 30

CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

arXiv:2607. 26752v1 Announce Type: new Abstract: Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning.

By Behraj Khan, Shabir Ahmad, Syed Ahmad Chan Bukhari, Tahir Qasim Syed
arXiv Computer Vision
Sep 25

BiCC: Bidirectional Connected-Component Loss for Instance-Aware Segmentation

The paper introduces BiCC, a bidirectional connected-component loss that pairs annotation- and prediction-derived partitions to score predicted components on their own scale. By deriving instances from predictions, BiCC directly penalizes false-positive components regardless of size, allowing a balance parameter to control the lesion-wise precision–recall trade-off. Across five datasets, BiCC outperforms existing instance-aware losses such as CC-DiceCE and blob loss in lesion-wise F1, and improves over DiceCE on multiple datasets.

By Luc Bouteille, Frederic Jonske, Jens Kleesiek, Alexander Jaus
arXiv Machine Learning
Sep 10

Conditional Validity for Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group

The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.

By Melika Baghi