The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.
By Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky
arXiv:2606. 11949v1 Announce Type: new Abstract: We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a classifier has moved out of distribution.
By Jun Wen Leong
SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.
By Liang You, Hengyu Shi, Dongwen Ou
arXiv:2609.11592v2 Announce Type: replace
Abstract: Transportation agencies increasingly predict crash-injury severity with statistical and machine-learning models, but these models do not state how...
By Amir Rafe, Subasish Das
arXiv:2606. 14238v1 Announce Type: cross Abstract: Safety certification of Vision-Language-Action (VLA) driving planners under ISO 21448 (SOTIF) rests on an Operational Design Domain (ODD) specification that answers two complementary questions: when does the planner start to fail, and how severely does it fail once it does?
By Abhinaw Priyadershi, Jelena Frtunikj
Conformal Calibration Transfer addresses the challenge of applying conformal prediction when labeled calibration data is only available in a source space, while predictions are needed in a target space linked via unlabeled paired observations. The proposed Transported Conformal Calibration (TCC) method transports source calibration into the target domain and then corrects residual mismatches using only unlabeled target inputs, with two variants: TCC‑KS, which conservatively adjusts calibration based on a label‑free uncertainty surrogate, and weighted‑TCC, which reweights transported calibration for efficiency when weights are stable. Finite‑sample target‑domain coverage guarantees are provided, and experiments on CIFAR‑100‑C, Tiny‑ImageNet‑C, and SEN12MS demonstrate reliable coverage transfer without labeled target data, along with label‑free diagnostics to signal when correction is required.
By Achref Doula
arXiv:2608. 05199v1 Announce Type: cross Abstract: Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique.
By Zhenpeng Li
arXiv:2606. 29403v1 Announce Type: cross Abstract: Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups.
By Louis Berthier, Ahmed Shokry, Maxime Moreaud, Guillaume Ramelet, Aymeric Dieuleveut
arXiv:2608.21262v1 Announce Type: cross
Abstract: Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cut...
By Adam Noonan
arXiv:2511.15146v2 Announce Type: replace
Abstract: Conformal prediction (CP) constructs uncertainty sets for model outputs with finite-sample coverage guarantees. Yet ranking scores is straightforwa...
By Eugene Ndiaye
Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups. We introduce Self-Organized Conformal Prediction (SOCP), a calibration scheme that discovers input-space groups with a Self-Organizing Map (SOM) and, at test time, draws a local calibration buffer from the query's best-matching unit (BMU) cell or a fixed grid neighborhood.
arXiv:2606. 14909v1 Announce Type: cross Abstract: We consider the problem of uncertainty quantification for a pretrained classification model deployed under unknown distribution shift.
By Yanfei Zhou, Rizal Fathony, Nam H. Nguyen, Matteo Sesia