arXiv Machine Learning By Fabio Centofanti, Mia Hubert, Peter J. Rousseeuw

Cellwise and Casewise Robust Covariance in High Dimensions

Read the original on arXiv Machine Learning →

arXiv:2505. 19925v2 Announce Type: replace-cross Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

Efficient Estimation of High Information Projections using Nearest Neighbours

The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.

By David P. Hofmeyr
arXiv Statistics ML
2d ago

Casewise and Cellwise Robust Tensor-on-Tensor Regression

arXiv:2603.25911v2 Announce Type: replace-cross Abstract: Tensor-on-tensor regression is an important tool for the analysis of tensor data, aiming to predict a set of response tensors from a correspo...

By Mehdi Hirari, Fabio Centofanti, Mia Hubert, Stefan Van Aelst
arXiv Statistics ML
Sep 18

Robust Multi-Task Learning for Principal Component Analysis

The paper introduces robust multi-task procedures for principal component analysis that leverage similarity across tasks to enhance eigenspace estimation while remaining resilient to outlier tasks. It establishes non-asymptotic convergence rates and demonstrates that the methods achieve minimax optimal performance across various regimes. One procedure, based on matrix-depth, attains optimal error dependence on the proportion of outlier tasks, addressing a key challenge in robust multi-task learning.

By Dali Liu, Haolei Weng
arXiv Statistics ML
Aug 31

Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions

The paper introduces a multivariate pseudo‑Voigt mixture model, combining Gaussian and Cauchy components with shared location and scale parameters, for robust clustering and outlier detection. Parameter estimation is performed using an EM algorithm that leverages latent variables for efficient likelihood inference. The authors evaluate the model through simulations and real data, comparing it to established robust mixtures such as contaminated normals, and demonstrate its effectiveness on heavy‑tailed datasets.

By Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
arXiv Machine Learning
Sep 3

Clustering Three-Way Data with Outliers

The paper introduces a method for clustering matrix-variate normal data that accounts for outliers. It extends the OCLUST algorithm by employing subset log-likelihood distributions and an iterative trimming procedure. This approach enables robust clustering of complex structured data such as images and time series.

By Katharine M. Clark, Paul D. McNicholas
arXiv Machine Learning
Jun 5

Anchor PCA

arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.

By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters