arXiv:2602. 13848v2 Announce Type: replace Abstract: We propose a sequential test for detecting arbitrary distribution shifts that allows conformal test martingales (CTMs) to work under a fixed, reference-conditional setting.
By Shalev Shaer, Yarin Bar, Drew Prinster, Yaniv Romano
arXiv:2608. 14089v1 Announce Type: new Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves.
By Thiago Sandoval, Ufuk Topcu
The paper introduces SHAP concentration as a pre‑deployment diagnostic for detecting when conformal prediction will fail under distribution shift, specifically in gradient‑boosted classifiers. Using a COVID‑19 supply‑chain case study, the authors show that higher feature‑importance concentration correlates with larger drops in coverage, while standard shift detectors cannot differentiate between catastrophic and robust outcomes. The diagnostic is validated on additional datasets, and a formal theorem links concentration to worsening conformity‑score bounds, though it does not capture global‑sensitivity failures in neural networks.
By Chorok Lee
arXiv:2606. 20859v2 Announce Type: replace-cross Abstract: A fundamental assumption in statistics and machine learning is that ``the future looks like the past,'' formalized as exchangeability: the joint data distribution is order-invariant.
By Johan Hallberg Szabadv\'ary
arXiv:2610.00873v1 Announce Type: cross
Abstract: In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between...
By Hongyu Cao, Xinyuan Wang, Arun Vignesh Malarkkan, Kunpeng Liu, Haifeng Chen, Yanjie Fu
arXiv:2608.30502v1 Announce Type: new
Abstract: Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is increasingly delegated to statistical moni...
By Weijia Han, Lisha Qu
arXiv:2606. 01256v1 Announce Type: cross Abstract: This paper introduces a distribution-free framework for constructing post-detection confidence sets for changepoints after stopping a sequential change detection procedure.
By Aytijhya Saha, Aaditya Ramdas
The paper introduces the Stream Cruise Control Method (SCCM), a framework for detecting and adapting to concept drift in online regression. SCCM performs early-response drift detection, quantifies drift magnitude, applies KPI-window-based thresholding to reduce false alarms, dynamically tunes hyperparameters, and recalibrates models, all within an in-memory design for real-time operation. Evaluations on synthetic and real-world datasets demonstrate that SCCM improves predictive performance compared to eight baseline detector–adaptation methods.
By Mohammad Abu-Shaira, Weishi Shi
arXiv:2609.26652v1 Announce Type: cross
Abstract: Commonly, classifiers and monitoring procedures are trained from labeled data by optimizing an objective such as the misclassification rate. This may...
By Ansgar Steland
arXiv:2606. 11949v1 Announce Type: new Abstract: We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a classifier has moved out of distribution.
By Jun Wen Leong
arXiv:2609.06173v1 Announce Type: new
Abstract: Post-deployment drift poses a critical risk to algorithmic accountability, particularly when ground truth labels are delayed and performance degradatio...
By Muhammad Rehman Zafar, Ali El-Sharif, Naimul Khan
Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition introduces C-PP-COAD, a framework that uses synthetic calibration data to reduce reliance on real-world calibration while maintaining assumption-free false discovery rate control. The method wraps any anomaly detection algorithm, converting its scores into conformal p-values for online testing. Experiments on synthetic and real datasets—including thyroid dysfunction, O‑RAN conflict, 5G intrusion, and UE throughput degradation—show that C-PP-COAD preserves FDR guarantees while significantly cutting the need for real calibration data.
By Amirmohammad Farzaneh, Osvaldo Simeone