arXiv Machine Learning

Monotonic anomaly detection

The paper introduces methods for monotonic anomaly detection, focusing on anomalies that exhibit high (or low) attribute values rather than arbitrary deviations. It proposes an asymmetrical distance measure using a ramp function for distance-based methods and a modified path length algorithm for Isolation Forest. Experiments on synthetic and real-life datasets demonstrate improved detection performance on datasets with monotonic attributes.

arXiv AI
Sep 25

Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data

The paper introduces a deep positive‑unlabeled anomaly detection framework that combines positive‑unlabeled learning with deep models such as autoencoders and deep support vector data descriptions. It addresses the issue of contaminated unlabeled data by approximating anomaly scores for normal data using both unlabeled and labeled anomaly samples, allowing training without labeled normal data. The authors provide a theoretical generalization error bound and demonstrate improved detection performance over existing methods on several datasets.

By Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai, Yuuki Yamanaka
arXiv Machine Learning
Jun 19

We Need to Rethink Benchmarking in Anomaly Detection

arXiv:2507. 15584v2 Announce Type: replace Abstract: Despite the continuous proposal of new anomaly detection algorithms and extensive benchmarking efforts, progress seems to stagnate, with only minor performance differences between established baselines and new algorithms.

By Philipp R\"ochner, Simon Kl\"uttermann, Kevin Kammler, Franz Rothlauf, Emmanuel M\"uller, Daniel Schl\"or
arXiv Machine Learning
Sep 15

Isolation-based Spherical Ensemble Representations for Tabular Anomaly Detection

The paper introduces ISER, an isolation-based method for unsupervised tabular anomaly detection that uses hypersphere radii to encode local density and maintains linear time and constant space complexity. ISER builds ensemble representations where smaller radii indicate dense regions and larger radii indicate sparse regions, and it employs a similarity-based scoring method that compares these representations to a theoretical anomaly reference pattern. Experiments on 20 real-world datasets show that ISER outperforms 12 state‑of‑the‑art methods, including an enhanced Isolation Forest.

By Yang Cao, Sikun Yang, Hao Tian, Kai He, Lianyong Qi, Ming Liu, Yujiu Yang, Hong-Kun Zhang
arXiv Machine Learning
Jul 31

ARES: Anomaly Recognition Model For Edge Streams

arXiv:2511. 22078v2 Announce Type: replace Abstract: Many real-world scenarios involving streaming information can be represented as temporal graphs, where data flows through dynamic changes in edges over time.

By Simone Mungari, Albert Bifet, Giuseppe Manco, Bernhard Pfahringer
arXiv Machine Learning
Sep 23

Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

The paper investigates whether the performance of anomaly detection systems can be predicted without labeled anomalies. For kNN-based detectors, it derives a lower bound on AUC that links detection performance to the separation and variance of inlier and outlier scores, and uses this to analyze how density variation, intrinsic dimensionality, and domain mismatch affect score variability. The authors introduce pseudo‑anomaly probes that provide a reference for estimating relative score separation, and demonstrate through experiments on DCASE benchmarks that these probes enable anomaly‑free model selection to outperform conventional development‑set selection, especially under domain shift.

By Kevin Wilkinghoff, Zheng-Hua Tan
arXiv Machine Learning
Sep 24

Anomaly-Free Self-Optimization via AUC Bounds

arXiv:2609.27362v1 Announce Type: new Abstract: Anomalies are rare, and anomalous data are often unavailable during development, making it difficult to determine which anomaly detection models and co...

By Kevin Wilkinghoff, Zheng-Hua Tan