When One Point Is Not Enough: Addressing Ambiguous Instances in Dimensionality Reduction by Splitting
arXiv:2605. 23540v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) methods are widely used to visualize high-dimensional data.
arXiv:2509. 04222v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) is widely used for visualizing high-dimensional data, often with the goal of revealing expected cluster structure.
arXiv:2605. 23540v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) methods are widely used to visualize high-dimensional data.
arXiv:2506. 08725v3 Announce Type: replace-cross Abstract: Misuses of t-SNE and UMAP in visual analytics have become increasingly common.
arXiv:2509. 03373v2 Announce Type: replace Abstract: Dimensionality reduction methods such as t-SNE and UMAP are popular methods for visualizing data with a potential (latent) clustered structure.
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.
arXiv:2609.26748v1 Announce Type: cross Abstract: Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters,...
arXiv:2606. 08258v1 Announce Type: cross Abstract: Understanding and comparing structures in scalar fields is a central challenge in scientific visualization, with applications ranging from feature analysis to temporal and structural comparison.
arXiv:2607. 27463v1 Announce Type: new Abstract: Dimensionality Reduction (DR) is a fundamental tool for high-dimensional data exploration, reducing the complexity of latent spaces of machine learning models, and assisting in the explanation of complex opaque models.
arXiv:2606. 00327v1 Announce Type: cross Abstract: Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries.
arXiv:2608. 06990v1 Announce Type: cross Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning.
The paper introduces the Rashomon set for dimension reduction, a collection of equally good embeddings that preserve high‑dimensional structure. It proposes PCA‑informed alignment to make axes interpretable, concept‑alignment regularization to incorporate external knowledge, and a method to extract trustworthy nearest‑neighbor relationships across the Rashomon set for refined embeddings. These techniques aim to produce interpretable, robust, and goal‑aligned visualizations by leveraging multiple valid embeddings instead of a single one.
arXiv:2607. 08746v1 Announce Type: cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally.