arXiv Machine Learning

Cellwise and Casewise Robust Covariance in High Dimensions

arXiv:2505. 19925v2 Announce Type: replace-cross Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers.

arXiv Machine Learning
Aug 27

Efficient Estimation of High Information Projections using Nearest Neighbours

The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.

By David P. Hofmeyr
arXiv Statistics ML
2d ago

Casewise and Cellwise Robust Tensor-on-Tensor Regression

arXiv:2603.25911v2 Announce Type: replace-cross Abstract: Tensor-on-tensor regression is an important tool for the analysis of tensor data, aiming to predict a set of response tensors from a correspo...

By Mehdi Hirari, Fabio Centofanti, Mia Hubert, Stefan Van Aelst
arXiv Statistics ML
Sep 18

Robust Multi-Task Learning for Principal Component Analysis

The paper introduces robust multi-task procedures for principal component analysis that leverage similarity across tasks to enhance eigenspace estimation while remaining resilient to outlier tasks. It establishes non-asymptotic convergence rates and demonstrates that the methods achieve minimax optimal performance across various regimes. One procedure, based on matrix-depth, attains optimal error dependence on the proportion of outlier tasks, addressing a key challenge in robust multi-task learning.

By Dali Liu, Haolei Weng
arXiv Statistics ML
Aug 31

Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions

The paper introduces a multivariate pseudo‑Voigt mixture model, combining Gaussian and Cauchy components with shared location and scale parameters, for robust clustering and outlier detection. Parameter estimation is performed using an EM algorithm that leverages latent variables for efficient likelihood inference. The authors evaluate the model through simulations and real data, comparing it to established robust mixtures such as contaminated normals, and demonstrate its effectiveness on heavy‑tailed datasets.

By Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
arXiv Machine Learning
Sep 3

Clustering Three-Way Data with Outliers

The paper introduces a method for clustering matrix-variate normal data that accounts for outliers. It extends the OCLUST algorithm by employing subset log-likelihood distributions and an iterative trimming procedure. This approach enables robust clustering of complex structured data such as images and time series.

By Katharine M. Clark, Paul D. McNicholas
arXiv Machine Learning
Jun 5

Anchor PCA

arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.

By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters
arXiv Machine Learning
4d ago

High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations

The paper investigates Partial Least Squares (PLS) in high-dimensional settings, focusing on a model where two data matrices share a low-rank latent structure plus individual-specific components. By analyzing the singular vectors of the cross‑covariance matrix with random matrix theory, the authors derive asymptotic characterizations of how well the estimated latent directions align with the true ones. They show that the PLS variant based on Singular Value Decomposition (PLS‑SVD) outperforms separate principal component analysis in detecting the common latent subspace, while also identifying regimes where PLS‑SVD behaves counter‑intuitively or reaches fundamental limits.

By Victor L\'eger, Florent Chatelain
arXiv Machine Learning
Jul 27

Heavy-Tailed Principal Component Analysis

arXiv:2603. 11308v3 Announce Type: replace Abstract: Principal Component Analysis (PCA) is a cornerstone of dimensionality reduction, yet its classical formulation relies critically on second-order moments and is therefore fragile in the presence of heavy-tailed data and impulsive noise.

By Mario Sayde, Christopher Khater, Jihad Fahs, Ibrahim Abou-Faycal
arXiv Machine Learning
Jun 16

Localized Kernel Projection Outlyingness: A Two-Stage Approach for Multi-Modal Outlier Detection

arXiv:2510. 24043v4 Announce Type: replace Abstract: This paper presents Two-Stage LKPLO, a novel multi-stage outlier detection framework that overcomes the coexisting limitations of conventional projection-based methods: their reliance on a fixed statistical metric and their assumption of a single data structure.

By Akira Tamamori
arXiv Machine Learning
Aug 20

Inference and Uncertainty Quantification for Streaming $r$-PCA

The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.

By Haoshu Xu, Hongzhe Li