arXiv Machine Learning

Robust Local Polynomial Regression with Similarity Kernels

arXiv:2501. 10729v3 Announce Type: replace-cross Abstract: Local Polynomial Regression (LPR) is a widely used nonparametric method for modeling complex relationships due to its flexibility and simplicity.

arXiv Machine Learning
Jun 16

Localized Kernel Projection Outlyingness: A Two-Stage Approach for Multi-Modal Outlier Detection

arXiv:2510. 24043v4 Announce Type: replace Abstract: This paper presents Two-Stage LKPLO, a novel multi-stage outlier detection framework that overcomes the coexisting limitations of conventional projection-based methods: their reliance on a fixed statistical metric and their assumption of a single data structure.

By Akira Tamamori
arXiv Machine Learning
Sep 16

A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure

The paper introduces Total Sensitivity Kernels (TSKs), a weighted ANOVA kernel framework that learns the importance of individual inputs and their interactions for approximating a multivariable black-box function from limited data. By selecting an RKHS where the target function has minimum norm, the authors derive a unique solution and prove consistency for finite-data interpolation. Numerical experiments show that adapting the kernel to the learned multivariable structure can significantly improve approximation accuracy compared to a standard product kernel.

By John E. Darges, Laura Weidensager
arXiv Statistics ML
Aug 31

Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions

The paper introduces a multivariate pseudo‑Voigt mixture model, combining Gaussian and Cauchy components with shared location and scale parameters, for robust clustering and outlier detection. Parameter estimation is performed using an EM algorithm that leverages latent variables for efficient likelihood inference. The authors evaluate the model through simulations and real data, comparing it to established robust mixtures such as contaminated normals, and demonstrate its effectiveness on heavy‑tailed datasets.

By Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
arXiv Machine Learning
Aug 6

Nonparametric Goodness-of-fit Testing under Covariate Shift

arXiv:2608. 04860v1 Announce Type: cross Abstract: This paper develops procedures for nonparametric goodness-of-fit testing under covariate shift, where labelled data are drawn from a source population but goodness-of-fit is evaluated for a target population.

By Zhen Hou, Dong Xia
Hugging Face Trending Papers
Aug 13

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data.