arXiv:2608. 11269v1 Announce Type: cross Abstract: Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging.
By Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry
arXiv:2407. 01718v2 Announce Type: replace-cross Abstract: Embedding high-dimensional data into a low-dimensional space is an indispensable component of data analysis.
By Boris Landa, Yuval Kluger, Rong Ma
arXiv:2608. 08704v1 Announce Type: cross Abstract: Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime.
By Zeqin Lin, Guangming Pan, Zhixiang Zhang, Yinbing Zhou
arXiv:2605. 23540v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) methods are widely used to visualize high-dimensional data.
By Diede P. M. van der Hoorn, Alessio Arleo, Fernando V. Paulovich
arXiv:2608. 05336v1 Announce Type: cross Abstract: Molecular representations are essential for the evaluation of molecular similarity and the development of structure-property relationships.
By Jacob W. Toney, Ayleen Y. Farnood, Samir Darouich, Heather J. Kulik
arXiv:2607. 19387v1 Announce Type: cross Abstract: Surrogate modeling for high-dimensional nonlinear dynamical systems that exhibit chaos requires mechanisms that preserve not only pointwise accuracy but also the scale-dependent structure of physical fields.
By Kanad Sen, Romit Maulik
arXiv:2607. 21039v1 Announce Type: new Abstract: Spectral methods are among the most widely used techniques for community detection, clustering, and graph learning.
By Zhuan Liang, Zheng Zhai
arXiv:2510. 02308v2 Announce Type: replace Abstract: Estimating the tangent spaces of a data manifold is a fundamental problem in geometric data analysis.
By Dhruv Kohli, Sawyer J. Robertson, Gal Mishne, Alexander Cloninger
The paper introduces an unsupervised framework that merges manifold learning with rank‑based interpretable graph embeddings to address the Geometric and Interpretability Gaps in visual representation learning. By first analyzing contextual information on the dataset manifold and then producing sparse, self‑explainable embeddings, the method achieves dimensionality reduction while preserving or improving performance in image retrieval and semi‑supervised Graph Convolutional Network classification. Experiments across varied datasets confirm that these context‑aware representations maintain high downstream effectiveness.
By Thiago C\'esar Castilho Almeida, Gustavo Rosseto Let\'icio, Vinicius Atsushi Sato Kawai, Daniel Carlos Guimar\~aes Pedronette
arXiv:2607. 08746v1 Announce Type: cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally.
By Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz
arXiv:2601.20173v3 Announce Type: replace
Abstract: We present a new nonlinear dimensionality reduction method, MAPLE, that enhances UMAP by improving manifold modeling. MAPLE employs a self-supervis...
By Zeyang Huang, Takanori Fujiwara, Angelos Chatzimparmpas, Wandrille Duchemin, Andreas Kerren
DMT‑Dens is a parametric manifold‑visualization technique that uses a latent‑token Transformer encoder to produce two‑dimensional embeddings of high‑dimensional biological data. It preserves sampling density by aligning rank‑based manifold structures and optimizing a Pearson‑correlation loss on k‑nearest‑neighbor log‑radius estimates. Benchmark tests show that DMT‑Dens maintains density fidelity while achieving competitive label separability on biological datasets.
By Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling, Fuji Yang, Zelin Zang