The paper introduces a multivariate pseudo‑Voigt mixture model, combining Gaussian and Cauchy components with shared location and scale parameters, for robust clustering and outlier detection. Parameter estimation is performed using an EM algorithm that leverages latent variables for efficient likelihood inference. The authors evaluate the model through simulations and real data, comparing it to established robust mixtures such as contaminated normals, and demonstrate its effectiveness on heavy‑tailed datasets.
By Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
arXiv:2508. 00110v2 Announce Type: replace-cross Abstract: Functional data present unique challenges for clustering due to their infinite-dimensional nature and potential sensitivity to outliers.
By Katharine M. Clark, Paul D. McNicholas
arXiv:2505. 19925v2 Announce Type: replace-cross Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers.
By Fabio Centofanti, Mia Hubert, Peter J. Rousseeuw
arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
By Naitik Gada (Rochester Institute of Technology)
arXiv:2608.30093v1 Announce Type: cross
Abstract: We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) me...
By Anirban Mondal, Paromita Banerjee, Abhijit Mandal
arXiv:2606. 19255v1 Announce Type: new Abstract: Time series anomaly detection plays a crucial role in a wide range of real-world applications.
By Xingze Zheng, Hanyin Cheng, Siyuan Wang, Yiting Hao, Peng Chen, Yuan Jun, Yang Shu
arXiv:2510. 24043v4 Announce Type: replace Abstract: This paper presents Two-Stage LKPLO, a novel multi-stage outlier detection framework that overcomes the coexisting limitations of conventional projection-based methods: their reliance on a fixed statistical metric and their assumption of a single data structure.
By Akira Tamamori
The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.
By David P. Hofmeyr
arXiv:2606. 18833v1 Announce Type: new Abstract: This paper introduces a semi-supervised clustering framework grounded in the statistical duality between grouping principles and anomaly detection.
By Nassir Mohammad
The paper extends the h‑plot multidimensional scaling technique to handle three‑way asymmetric dissimilarities, enabling the extraction of archetypal profiles and clustering in a unified Euclidean space. It provides an explicit eigenvector‑based solution that avoids local minima, is scale‑invariant, computationally efficient, and includes a straightforward goodness‑of‑fit assessment. The method is benchmarked against existing models and demonstrated on a financial dataset, with all data and code made publicly available for reproducibility.
By Mireia Mollar-Gumbau, Aleix Alcacer, Rafael Benitez, Vicente J. Bolos, Irene Epifanio
arXiv:2608. 11375v1 Announce Type: new Abstract: In this paper, we test different methods for multivariate detection of outliers in the data of transmitted heat energy in the selected substation of local District Heating System, by also considering outside ambient temperature, namely Z-score (univariate, as a benchmark), Mahalanobis distances, Principal Component Analysis (PCA), Isolation Forest and Hotelling's T-squared test.
By Rajko Turudija, Du\v{s}an Stojiljkovi\'c, Milan Zdravkovi\'c, Marko Ignjatovi\'c
arXiv:2607. 24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge.
By Filip Kosiorowski, Grzegorz Sroka