arXiv:2510. 25599v2 Announce Type: replace Abstract: Regression tasks, notably in safety-critical domains, require reliable uncertainty quantification, yet the literature remains largely classification-focused.
By Christopher B\"ulte, Yusuf Sale, Gitta Kutyniok, Eyke H\"ullermeier
arXiv:2609. 21342v1 Announce Type: cross Abstract: Real data often contain unusual observations that can exert disproportionate effects on variable selection, especially in complex predictor settings.
By Abdul-Nasah Soale, Adewale F. Lukman, Essoham Ali
arXiv:2411.02771v3 Announce Type: replace-cross
Abstract: Doubly robust estimators are widely used for estimating average treatment effects and other linear summaries of regression functions. While c...
By Lars van der Laan, Alex Luedtke, Marco Carone
arXiv:2510. 24043v4 Announce Type: replace Abstract: This paper presents Two-Stage LKPLO, a novel multi-stage outlier detection framework that overcomes the coexisting limitations of conventional projection-based methods: their reliance on a fixed statistical metric and their assumption of a single data structure.
By Akira Tamamori
arXiv:2602. 13362v2 Announce Type: replace-cross Abstract: A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty.
By \'Ad\'am Jung, Domokos M. Kelen, Andr\'as A. Bencz\'ur
The paper introduces Total Sensitivity Kernels (TSKs), a weighted ANOVA kernel framework that learns the importance of individual inputs and their interactions for approximating a multivariable black-box function from limited data. By selecting an RKHS where the target function has minimum norm, the authors derive a unique solution and prove consistency for finite-data interpolation. Numerical experiments show that adapting the kernel to the learned multivariable structure can significantly improve approximation accuracy compared to a standard product kernel.
By John E. Darges, Laura Weidensager
arXiv:2609.15785v1 Announce Type: cross
Abstract: We study density ratio estimation and importance-weighted regression under target shift with continuous outputs. Under target shift, the conditional...
By Ren-Rui Liu, Zheng-Chu Guo
arXiv:2608. 13418v1 Announce Type: cross Abstract: Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution.
By Yikai Xu, Zhao Chen, Jian Huang
The paper introduces a multivariate pseudo‑Voigt mixture model, combining Gaussian and Cauchy components with shared location and scale parameters, for robust clustering and outlier detection. Parameter estimation is performed using an EM algorithm that leverages latent variables for efficient likelihood inference. The authors evaluate the model through simulations and real data, comparing it to established robust mixtures such as contaminated normals, and demonstrate its effectiveness on heavy‑tailed datasets.
By Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
arXiv:2608. 04860v1 Announce Type: cross Abstract: This paper develops procedures for nonparametric goodness-of-fit testing under covariate shift, where labelled data are drawn from a source population but goodness-of-fit is evaluated for a target population.
By Zhen Hou, Dong Xia
arXiv:2608. 13590v1 Announce Type: new Abstract: XGBoost is a very popular and powerful method for prediction.
By Iris Arag\'on Mladosich, Christophe Croux
Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data.