On Out-of-sample Embedding in UMAP
arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.
arXiv:2607. 08746v1 Announce Type: cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally.
arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.
arXiv:2509. 03373v2 Announce Type: replace Abstract: Dimensionality reduction methods such as t-SNE and UMAP are popular methods for visualizing data with a potential (latent) clustered structure.
arXiv:2605. 23540v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) methods are widely used to visualize high-dimensional data.
arXiv:2608. 06990v1 Announce Type: cross Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning.
arXiv:2604.02535v2 Announce Type: replace Abstract: Dimensionality reduction (DR) involves two longstanding trade-offs. First, preserving local neighborhoods can come at the cost of global structure....
arXiv:2507.23559v2 Announce Type: replace-cross Abstract: Certain data are naturally modeled by networks or weighted graphs, be they biological networks or mobility networks. When there is no canonic...
arXiv:2608. 11269v1 Announce Type: cross Abstract: Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging.
arXiv:2608.29001v1 Announce Type: new Abstract: In a data-driven world, efficiently organizing and mapping relationships between objects is crucial. Graphs are powerful tools for modeling these conne...
The paper introduces CluProp, a framework that treats varied‑density clustering in high‑dimensional spaces as a label propagation process over neighborhood graphs. By combining density‑based ideas with graph connectivity, it offers a deterministic propagation strategy that reduces parameter sensitivity and scales efficiently to millions of points. CluProp is agnostic to distance metrics and consistently outperforms existing baselines in accuracy while processing large datasets in minutes.
arXiv:2508. 02989v2 Announce Type: replace Abstract: We propose a novel perspective on varied-density clustering for high-dimensional data by framing it as a label propagation process in neighborhood graphs that adapt to local density variations.
arXiv:2203. 04711v2 Announce Type: replace Abstract: We present a framework for embedding graph structured data into a vector space, taking into account node features and topology of a graph into the optimal transport (OT) problem.
The paper introduces an unsupervised framework that merges manifold learning with rank‑based interpretable graph embeddings to address the Geometric and Interpretability Gaps in visual representation learning. By first analyzing contextual information on the dataset manifold and then producing sparse, self‑explainable embeddings, the method achieves dimensionality reduction while preserving or improving performance in image retrieval and semi‑supervised Graph Convolutional Network classification. Experiments across varied datasets confirm that these context‑aware representations maintain high downstream effectiveness.