arXiv Statistics ML

Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport

arXiv Machine Learning
Sep 24

CAST: Context- and Anomaly Structure-Conditioned Time Series Anomaly Generation

CAST is a framework for generating anomalous time series that addresses the scarcity and heterogeneity of anomaly data. It uses a two‑stage approach: pretraining on abundant normal data to learn system dynamics, then finetuning with anomaly structure representations to capture diverse anomaly morphologies. Experiments on real‑world datasets show that CAST outperforms existing methods in both generation quality and downstream task performance.

By Haochen Zhang, Jie Peng, Songyuan Sui, Yu-Chao Huang, Xiangqi Zhu, Tianlong Chen
Hugging Face Trending Papers
Aug 13

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data.

arXiv Machine Learning
Sep 11

Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning

The paper introduces a framework for out-of-distribution (OOD) detection that addresses the trade‑off between detection performance and classification accuracy caused by fine‑tuning with auxiliary outlier data. It optimizes three factors—model reminder, data sampling, and representation learning—by proposing Self‑Knowledge Distillation to preserve accuracy, Semi‑hard Outlier Sampling to enhance detection with minimal data, and Outlier‑aware Supervised Contrastive Learning to improve ID‑OOD separability. The combined approach yields cumulative gains, outperforming existing methods on diverse benchmarks, especially in long‑tailed scenarios, and offers a robust baseline for real‑world OOD detection.

By Hyunjun Choi, JaeHo Chung, Hawook Jeong
arXiv Machine Learning
Sep 23

Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

The paper investigates whether the performance of anomaly detection systems can be predicted without labeled anomalies. For kNN-based detectors, it derives a lower bound on AUC that links detection performance to the separation and variance of inlier and outlier scores, and uses this to analyze how density variation, intrinsic dimensionality, and domain mismatch affect score variability. The authors introduce pseudo‑anomaly probes that provide a reference for estimating relative score separation, and demonstrate through experiments on DCASE benchmarks that these probes enable anomaly‑free model selection to outperform conventional development‑set selection, especially under domain shift.

By Kevin Wilkinghoff, Zheng-Hua Tan
arXiv AI
Jun 4

TPA-AD: A Two-Stage Pseudo Anomaly-Guided Method for Bearing Time-Series Anomaly Detection

arXiv:2606. 04073v1 Announce Type: cross Abstract: This paper proposes a two-stage pseudo anomaly-guided anomaly detection method (\textbf{T}wo-stage \textbf{P}seudo \textbf{A}nomaly-guided \textbf{A}nomaly \textbf{D}etection, \textbf{TPA-AD}) for axle-box bearing time-series anomaly detection (time series anomaly detection, TSAD) under the setting where only normal samples are available for training.

By Xiancheng Wang, Zhibo Zhang, Ran Li, Rui Wang, Minghang Zhao, Shisheng Zhong, Lin Wang