arXiv AI

WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales

arXiv:2505. 04608v5 Announce Type: replace-cross Abstract: Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof of system reliability, but also continual, post-deployment monitoring to quickly detect and address any unsafe behavior.

arXiv AI
Aug 28

Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study

The paper introduces SHAP concentration as a pre‑deployment diagnostic for detecting when conformal prediction will fail under distribution shift, specifically in gradient‑boosted classifiers. Using a COVID‑19 supply‑chain case study, the authors show that higher feature‑importance concentration correlates with larger drops in coverage, while standard shift detectors cannot differentiate between catastrophic and robust outcomes. The diagnostic is validated on additional datasets, and a formal theorem links concentration to worsening conformity‑score bounds, though it does not capture global‑sensitivity failures in neural networks.

By Chorok Lee
arXiv AI
Sep 11

SCCM : Stream Cruise Control Method for Automated Drift Detection and Adaptation

The paper introduces the Stream Cruise Control Method (SCCM), a framework for detecting and adapting to concept drift in online regression. SCCM performs early-response drift detection, quantifies drift magnitude, applies KPI-window-based thresholding to reduce false alarms, dynamically tunes hyperparameters, and recalibrates models, all within an in-memory design for real-time operation. Evaluations on synthetic and real-world datasets demonstrate that SCCM improves predictive performance compared to eight baseline detector–adaptation methods.

By Mohammad Abu-Shaira, Weishi Shi
arXiv Machine Learning
Aug 20

Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition

Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition introduces C-PP-COAD, a framework that uses synthetic calibration data to reduce reliance on real-world calibration while maintaining assumption-free false discovery rate control. The method wraps any anomaly detection algorithm, converting its scores into conformal p-values for online testing. Experiments on synthetic and real datasets—including thyroid dysfunction, O‑RAN conflict, 5G intrusion, and UE throughput degradation—show that C-PP-COAD preserves FDR guarantees while significantly cutting the need for real calibration data.

By Amirmohammad Farzaneh, Osvaldo Simeone