arXiv Machine Learning

Early Failure Prediction from Near-Anomaly Detection: A Proactive Approach

arXiv:2607. 26704v1 Announce Type: cross Abstract: Anomaly detection methods often have uncertain behavior with respect to samples near the distribution boundary, limiting their ability to anticipate future anomalies.

arXiv Machine Learning
Aug 20

Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition

Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition introduces C-PP-COAD, a framework that uses synthetic calibration data to reduce reliance on real-world calibration while maintaining assumption-free false discovery rate control. The method wraps any anomaly detection algorithm, converting its scores into conformal p-values for online testing. Experiments on synthetic and real datasets—including thyroid dysfunction, O‑RAN conflict, 5G intrusion, and UE throughput degradation—show that C-PP-COAD preserves FDR guarantees while significantly cutting the need for real calibration data.

By Amirmohammad Farzaneh, Osvaldo Simeone
arXiv Machine Learning
Sep 24

CAST: Context- and Anomaly Structure-Conditioned Time Series Anomaly Generation

CAST is a framework for generating anomalous time series that addresses the scarcity and heterogeneity of anomaly data. It uses a two‑stage approach: pretraining on abundant normal data to learn system dynamics, then finetuning with anomaly structure representations to capture diverse anomaly morphologies. Experiments on real‑world datasets show that CAST outperforms existing methods in both generation quality and downstream task performance.

By Haochen Zhang, Jie Peng, Songyuan Sui, Yu-Chao Huang, Xiangqi Zhu, Tianlong Chen
arXiv Machine Learning
Sep 23

Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

The paper investigates whether the performance of anomaly detection systems can be predicted without labeled anomalies. For kNN-based detectors, it derives a lower bound on AUC that links detection performance to the separation and variance of inlier and outlier scores, and uses this to analyze how density variation, intrinsic dimensionality, and domain mismatch affect score variability. The authors introduce pseudo‑anomaly probes that provide a reference for estimating relative score separation, and demonstrate through experiments on DCASE benchmarks that these probes enable anomaly‑free model selection to outperform conventional development‑set selection, especially under domain shift.

By Kevin Wilkinghoff, Zheng-Hua Tan
arXiv AI
Aug 19

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

The paper introduces LoRD, a lightweight post‑hoc calibration framework designed to improve confidence reliability in language‑model‑based log anomaly detectors. LoRD learns route‑specific reliability models from latent representations of correctly classified validation samples and uses reconstruction distances to estimate prediction reliability. By selectively recalibrating high‑risk predictions, LoRD reduces overconfident errors while maintaining strong anomaly detection performance across four large‑scale log benchmark datasets.

By Bin Li, Dongdong Wang, Siyang Lu