arXiv:2609.38263v1 Announce Type: new
Abstract: Feature selection in neural networks remains a challenging problem, particularly in the presence of noisy or contaminated data. LassoNet is a recent ap...
By Daniela De Canditiis, Italia De Feis, Paola Stolfi
arXiv:2608. 13418v1 Announce Type: cross Abstract: Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution.
By Yikai Xu, Zhao Chen, Jian Huang
Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data.
arXiv:2606. 27385v1 Announce Type: new Abstract: The most widely used RANSAC variants score candidate models by counting inliers or summing per-point scores that saturate beyond a residual threshold.
By James Pritts, Felix Seegr\"aber, Kevin K\"oser
arXiv:2510. 24043v4 Announce Type: replace Abstract: This paper presents Two-Stage LKPLO, a novel multi-stage outlier detection framework that overcomes the coexisting limitations of conventional projection-based methods: their reliance on a fixed statistical metric and their assumption of a single data structure.
By Akira Tamamori
arXiv:2606. 27685v1 Announce Type: cross Abstract: Pervasive data contamination -- stemming from measurement errors, outliers, or adversarial corruption -- has motivated the development of robust statistical methods.
By Shixiang Liu, Hanming Yang
The paper investigates whether the performance of anomaly detection systems can be predicted without labeled anomalies. For kNN-based detectors, it derives a lower bound on AUC that links detection performance to the separation and variance of inlier and outlier scores, and uses this to analyze how density variation, intrinsic dimensionality, and domain mismatch affect score variability. The authors introduce pseudo‑anomaly probes that provide a reference for estimating relative score separation, and demonstrate through experiments on DCASE benchmarks that these probes enable anomaly‑free model selection to outperform conventional development‑set selection, especially under domain shift.
By Kevin Wilkinghoff, Zheng-Hua Tan
arXiv:2608.22597v1 Announce Type: new
Abstract: Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect est...
By Jing Wang, HaiYing Wang, Qiang Zhang, Hao Helen Zhang
arXiv:2603.25911v2 Announce Type: replace-cross
Abstract: Tensor-on-tensor regression is an important tool for the analysis of tensor data, aiming to predict a set of response tensors from a correspo...
By Mehdi Hirari, Fabio Centofanti, Mia Hubert, Stefan Van Aelst
arXiv:2510. 06505v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection plays a crucial role in ensuring the robustness of machine learning systems deployed in real-world applications.
By Momin Abbas, Ali Falahati, Hossein Goli, Mohammad Mohammadi Amiri
arXiv:2505. 19925v2 Announce Type: replace-cross Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers.
By Fabio Centofanti, Mia Hubert, Peter J. Rousseeuw
arXiv:2607. 03839v1 Announce Type: new Abstract: Sparse feature selection is critical for high-dimensional machine learning, yet traditional $\ell_1$-regularized methods are often brittle under observational noise and spurious correlations, leading to unstable feature supports and degraded generalization.
By Zhen Huang, Peicheng Xu, Junbiao Pang, Yulong Zheng