arXiv Machine Learning

CS-WCP: Robust Conformal Sets for LLM-Judge Traffic Shifts with Uncertain Group Proportions

CS-WCP introduces confidence‑set weighted conformal prediction to provide robust prediction sets for large‑language‑model judges when deployment traffic shifts the prevalence of task or policy groups. By constructing simultaneous exact intervals for source and target group masses and taking the union over all compatible ratio vectors, CS‑WCP achieves high coverage (mean 0.973) with few failures across 336 constructed traffic shifts, outperforming standard source conformal prediction. The method offers an auditable coverage safeguard under uncertain mixture weights, focusing on conservative tail protection rather than tighter set sizes.

arXiv Machine Learning
Sep 21

Available Guardrails: Certifying Selective Prediction across ML Systems

The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.

By Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky
arXiv Machine Learning
1d ago

SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions

SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.

By Liang You, Hengyu Shi, Dongwen Ou
arXiv Machine Learning
Sep 11

Conformal Calibration Transfer

Conformal Calibration Transfer addresses the challenge of applying conformal prediction when labeled calibration data is only available in a source space, while predictions are needed in a target space linked via unlabeled paired observations. The proposed Transported Conformal Calibration (TCC) method transports source calibration into the target domain and then corrects residual mismatches using only unlabeled target inputs, with two variants: TCC‑KS, which conservatively adjusts calibration based on a label‑free uncertainty surrogate, and weighted‑TCC, which reweights transported calibration for efficiency when weights are stable. Finite‑sample target‑domain coverage guarantees are provided, and experiments on CIFAR‑100‑C, Tiny‑ImageNet‑C, and SEN12MS demonstrate reliable coverage transfer without labeled target data, along with label‑free diagnostics to signal when correction is required.

By Achref Doula
Hugging Face Trending Papers
Jun 28

Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery

Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups. We introduce Self-Organized Conformal Prediction (SOCP), a calibration scheme that discovers input-space groups with a Self-Organizing Map (SOM) and, at test time, draws a local calibration buffer from the query's best-matching unit (BMU) cell or a fixed grid neighborhood.