arXiv:2509. 04222v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) is widely used for visualizing high-dimensional data, often with the goal of revealing expected cluster structure.
By Diede P. M. van der Hoorn, Alessio Arleo, Fernando V. Paulovich
The paper introduces the Rashomon set for dimension reduction, a collection of equally good embeddings that preserve high‑dimensional structure. It proposes PCA‑informed alignment to make axes interpretable, concept‑alignment regularization to incorporate external knowledge, and a method to extract trustworthy nearest‑neighbor relationships across the Rashomon set for refined embeddings. These techniques aim to produce interpretable, robust, and goal‑aligned visualizations by leveraging multiple valid embeddings instead of a single one.
By Yiyang Sun, Haiyang Huang, Gaurav Rajesh Parikh, Cynthia Rudin
arXiv:2509. 03373v2 Announce Type: replace Abstract: Dimensionality reduction methods such as t-SNE and UMAP are popular methods for visualizing data with a potential (latent) clustered structure.
By Elizabeth Coda, Ery Arias-Castro, Gal Mishne
arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.
By Mohammad Tariqul Islam, Jason W. Fleischer
arXiv:2607. 27463v1 Announce Type: new Abstract: Dimensionality Reduction (DR) is a fundamental tool for high-dimensional data exploration, reducing the complexity of latent spaces of machine learning models, and assisting in the explanation of complex opaque models.
By Lucas Greff Meneses, Evandro S. Ortigossa, Claudio Silva, Luis Gustavo Nonato
arXiv:2608. 11269v1 Announce Type: cross Abstract: Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging.
By Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry
arXiv:2604.02535v2 Announce Type: replace
Abstract: Dimensionality reduction (DR) involves two longstanding trade-offs. First, preserving local neighborhoods can come at the cost of global structure....
By Zeyang Huang, Angelos Chatzimparmpas, Thomas H\"ollt, Takanori Fujiwara
arXiv:2607. 28324v1 Announce Type: new Abstract: Quality metrics play a crucial role in the proper use of dimensionality reduction projections for visual analysis of high-dimensional data.
By Jaume Ros, Alessio Arleo, Fernando Paulovich
arXiv:2607. 08746v1 Announce Type: cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally.
By Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz
arXiv:2606. 31119v1 Announce Type: new Abstract: Graphs are commonly visualized in 2D, where humans readily interpret spatial relationships, yet such layouts often distort higher-dimensional structure.
By Ya Ji (Khoury College of Computer Sciences, Northeastern University, Seattle), Xuefeng Li (Khoury College of Computer Sciences, Northeastern University, Seattle), Timo Brand (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Jacob Miller (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Peng Zhang (Khoury College of Computer Sciences, Northeastern University, Seattle), Stephen Kobourov (School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany), Yifan Hu (Khoury College of Computer Sciences, Northeastern University, Seattle)
arXiv:2606. 08258v1 Announce Type: cross Abstract: Understanding and comparing structures in scalar fields is a central challenge in scientific visualization, with applications ranging from feature analysis to temporal and structural comparison.
By Guangyu Meng, Mingzhe Li, Erin Wolf Chambers
3D vision-language models (3D VLMs) enable spatial reasoning over multi-view scenes but suffer from substantial token redundancy due to duplicated observations and large uninformative regions, leading to high computational cost. Although visual token compression has shown promise in accelerating 2D VLMs, it fails to capture the structured nature of 3D scenes and leads to incomplete spatial coverage and loss of fine-grained details.