arXiv:2606. 18833v1 Announce Type: new Abstract: This paper introduces a semi-supervised clustering framework grounded in the statistical duality between grouping principles and anomaly detection.
By Nassir Mohammad
arXiv:2606. 28970v1 Announce Type: cross Abstract: Unsupervised tabular anomaly detection requires methods that are accurate, robust across heterogeneous datasets, and computationally efficient.
By Quanling Zhao, Jiaying Yang, Ye Tian, Josh Victoria, Zhijun Wang, Pietro Mercati, Onat Gungor, Tajana Rosing
The paper investigates whether the performance of anomaly detection systems can be predicted without labeled anomalies. For kNN-based detectors, it derives a lower bound on AUC that links detection performance to the separation and variance of inlier and outlier scores, and uses this to analyze how density variation, intrinsic dimensionality, and domain mismatch affect score variability. The authors introduce pseudo‑anomaly probes that provide a reference for estimating relative score separation, and demonstrate through experiments on DCASE benchmarks that these probes enable anomaly‑free model selection to outperform conventional development‑set selection, especially under domain shift.
By Kevin Wilkinghoff, Zheng-Hua Tan
The paper introduces a deep positive‑unlabeled anomaly detection framework that combines positive‑unlabeled learning with deep models such as autoencoders and deep support vector data descriptions. It addresses the issue of contaminated unlabeled data by approximating anomaly scores for normal data using both unlabeled and labeled anomaly samples, allowing training without labeled normal data. The authors provide a theoretical generalization error bound and demonstrate improved detection performance over existing methods on several datasets.
By Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai, Yuuki Yamanaka
The paper introduces methods for monotonic anomaly detection, focusing on anomalies that exhibit high (or low) attribute values rather than arbitrary deviations. It proposes an asymmetrical distance measure using a ramp function for distance-based methods and a modified path length algorithm for Isolation Forest. Experiments on synthetic and real-life datasets demonstrate improved detection performance on datasets with monotonic attributes.
By Oliver Urs Lenz, Matthijs van Leeuwen
arXiv:2606. 15280v1 Announce Type: new Abstract: Most existing anomaly detection methods rely on estimating a probability density or learning an enclosing decision boundary, implicitly assuming that normal data occupies a region of non-zero volume in the ambient space.
By Alexander Bauer