arXiv Machine Learning

Clustering Three-Way Data with Outliers

The paper introduces a method for clustering matrix-variate normal data that accounts for outliers. It extends the OCLUST algorithm by employing subset log-likelihood distributions and an iterative trimming procedure. This approach enables robust clustering of complex structured data such as images and time series.

arXiv Statistics ML
Aug 31

Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions

The paper introduces a multivariate pseudo‑Voigt mixture model, combining Gaussian and Cauchy components with shared location and scale parameters, for robust clustering and outlier detection. Parameter estimation is performed using an EM algorithm that leverages latent variables for efficient likelihood inference. The authors evaluate the model through simulations and real data, comparing it to established robust mixtures such as contaminated normals, and demonstrate its effectiveness on heavy‑tailed datasets.

By Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
arXiv Machine Learning
Jun 16

Localized Kernel Projection Outlyingness: A Two-Stage Approach for Multi-Modal Outlier Detection

arXiv:2510. 24043v4 Announce Type: replace Abstract: This paper presents Two-Stage LKPLO, a novel multi-stage outlier detection framework that overcomes the coexisting limitations of conventional projection-based methods: their reliance on a fixed statistical metric and their assumption of a single data structure.

By Akira Tamamori
arXiv Machine Learning
Aug 27

Efficient Estimation of High Information Projections using Nearest Neighbours

The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.

By David P. Hofmeyr
arXiv Statistics ML
Sep 3

Multidimensional scaling of two-mode three-way asymmetric dissimilarities: finding archetypal profiles and clustering

The paper extends the h‑plot multidimensional scaling technique to handle three‑way asymmetric dissimilarities, enabling the extraction of archetypal profiles and clustering in a unified Euclidean space. It provides an explicit eigenvector‑based solution that avoids local minima, is scale‑invariant, computationally efficient, and includes a straightforward goodness‑of‑fit assessment. The method is benchmarked against existing models and demonstrated on a financial dataset, with all data and code made publicly available for reproducibility.

By Mireia Mollar-Gumbau, Aleix Alcacer, Rafael Benitez, Vicente J. Bolos, Irene Epifanio
arXiv Machine Learning
Aug 13

Towards an approach to multivariate outlier detection for District Heating System data

arXiv:2608. 11375v1 Announce Type: new Abstract: In this paper, we test different methods for multivariate detection of outliers in the data of transmitted heat energy in the selected substation of local District Heating System, by also considering outside ambient temperature, namely Z-score (univariate, as a benchmark), Mahalanobis distances, Principal Component Analysis (PCA), Isolation Forest and Hotelling's T-squared test.

By Rajko Turudija, Du\v{s}an Stojiljkovi\'c, Milan Zdravkovi\'c, Marko Ignjatovi\'c