arXiv:2604.02535v2 Announce Type: replace
Abstract: Dimensionality reduction (DR) involves two longstanding trade-offs. First, preserving local neighborhoods can come at the cost of global structure....
By Zeyang Huang, Angelos Chatzimparmpas, Thomas H\"ollt, Takanori Fujiwara
arXiv:2608. 14951v1 Announce Type: new Abstract: Low-rank matrix decompositions can uncover patterns and structure in data and have a number of different applications across many disciplines.
By Ying-Qiu Zheng, Alex Fung, Stephen M Smith, Rogier B Mars, Saad Jbabdi
arXiv:2607. 05464v1 Announce Type: cross Abstract: The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects.
By Yiqun Zhang, Yiu-ming Cheung
arXiv:2608. 06990v1 Announce Type: cross Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning.
By Yuning Yu, Jos\'e Rodr\'iguez-Pi\~neiro, Xuefeng Yin, Bin Feng
The paper introduces View distance, a novel metric that projects high‑dimensional data onto all pairwise two‑dimensional planes and sums the Euclidean distances across these projections. It satisfies metric axioms, couples features, suppresses redundancy, and captures anisotropic geometry. To make it scalable, the authors propose a plane‑selection strategy using iterative Maximum Weight Matching, reducing complexity from ω(n²) to ω(k) and demonstrating competitive performance on twelve datasets.
By Yiqun Zhang, Hou-biao Li
arXiv:2608. 11269v1 Announce Type: cross Abstract: Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging.
By Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry
arXiv:2607. 08746v1 Announce Type: cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally.
By Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz
The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.
By David P. Hofmeyr
This paper examines how metric adjustments to Multidimensional Scaling (MDS) can enhance its effectiveness as a visual tool for pattern recognition. The distance under consideration, referred to as Max-D-SW, is an adjustment of the Max-Sliced Wasserstein distance.
arXiv:2607. 15018v1 Announce Type: cross Abstract: High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables.
By Chun-houh Chen, Shun-Chuan Chang, Chiun-How Kao, Yi-Ju Lee, Shang-Ying Shiu, Yin-Jing Tien, ShengLi Tzeng, Han-Ming Wu
arXiv:2608.28799v1 Announce Type: cross
Abstract: Separable nonnegative matrix factorization (SNMF) has been widely used for low-rank representation and clustering of nonnegative data, owing to its a...
By Matthew McCarver, Jing Qin
The paper introduces a method for clustering matrix-variate normal data that accounts for outliers. It extends the OCLUST algorithm by employing subset log-likelihood distributions and an iterative trimming procedure. This approach enables robust clustering of complex structured data such as images and time series.
By Katharine M. Clark, Paul D. McNicholas