The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.
By Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky
arXiv:2606. 11949v1 Announce Type: new Abstract: We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a classifier has moved out of distribution.
By Jun Wen Leong
SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.
By Liang You, Hengyu Shi, Dongwen Ou
arXiv:2609.11592v2 Announce Type: replace
Abstract: Transportation agencies increasingly predict crash-injury severity with statistical and machine-learning models, but these models do not state how...
By Amir Rafe, Subasish Das
arXiv:2606. 14238v1 Announce Type: cross Abstract: Safety certification of Vision-Language-Action (VLA) driving planners under ISO 21448 (SOTIF) rests on an Operational Design Domain (ODD) specification that answers two complementary questions: when does the planner start to fail, and how severely does it fail once it does?
By Abhinaw Priyadershi, Jelena Frtunikj
Conformal Calibration Transfer addresses the challenge of applying conformal prediction when labeled calibration data is only available in a source space, while predictions are needed in a target space linked via unlabeled paired observations. The proposed Transported Conformal Calibration (TCC) method transports source calibration into the target domain and then corrects residual mismatches using only unlabeled target inputs, with two variants: TCC‑KS, which conservatively adjusts calibration based on a label‑free uncertainty surrogate, and weighted‑TCC, which reweights transported calibration for efficiency when weights are stable. Finite‑sample target‑domain coverage guarantees are provided, and experiments on CIFAR‑100‑C, Tiny‑ImageNet‑C, and SEN12MS demonstrate reliable coverage transfer without labeled target data, along with label‑free diagnostics to signal when correction is required.
By Achref Doula