arXiv:2602. 13848v2 Announce Type: replace Abstract: We propose a sequential test for detecting arbitrary distribution shifts that allows conformal test martingales (CTMs) to work under a fixed, reference-conditional setting.
By Shalev Shaer, Yarin Bar, Drew Prinster, Yaniv Romano
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
By Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth, Fran\c{c}ois Caron
arXiv:2505. 04608v5 Announce Type: replace-cross Abstract: Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof of system reliability, but also continual, post-deployment monitoring to quickly detect and address any unsafe behavior.
By Drew Prinster, Xing Han, Anqi Liu, Suchi Saria
arXiv:2607. 22985v1 Announce Type: cross Abstract: Conformalized selection has been widely applied to select high-quality candidates from large datasets with rigorous uncertainty quantification, such as reliable labeling, drug discovery, and the alignment of large language models.
By Chengyao Yu, Hongxin Wei, Bingyi Jing
arXiv:2609. 11235v1 Announce Type: new Abstract: Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse.
By Kartik Jhawar, Lipo Wang
The paper introduces the Label-Shift-Adjusted Bayesian Score (LSA score), a nonconformity measure for conformal prediction that corrects Bayesian scores under label shift by applying an importance-weighted transformation of the source predictive distribution. Unlike residual-based scores that produce uniform-width intervals, the LSA score yields shorter, adaptive intervals while maintaining comparable coverage in the target domain. Experiments on molecular property prediction demonstrate that the LSA score outperforms both residual-based and source-based Bayesian scores, though all methods experience some coverage loss under stronger shifts due to density-ratio estimation challenges.
By Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, Hyunjin Shin