arXiv Machine Learning By Hongmin Li

FastUMAP: Scalable Dimensionality Reduction via Bipartite Landmark Sampling

Read the original on arXiv Machine Learning →

arXiv:2605. 11428v2 Announce Type: replace Abstract: Exploratory analysis of high-dimensional data rarely stops at a single embedding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs

The paper studies Online Kernel Supervised Principal Component Analysis (OKSPCA), which uses random features and an Adam-style orthonormal basis update to optimize a supervised spectral objective. It shows that accurate optimization of this objective does not guarantee accurate population subspace recovery or improved predictive performance, and it provides theoretical results on consistency, concentration, and perturbation of the estimator. Empirical experiments on six benchmarks reveal that replacing the tracker with the exact empirical target does not significantly change regression deficits, while classification-rank models capture most of the terminal objective energy but can exhibit substantial geometric deviation; sample-size studies further separate empirical accuracy from population recovery. The diagnostics also compare computational trade-offs, indicating that exact on-request computation can be faster in classification settings, whereas Adam saves time relative to full thin‑SVD in some dense regression requests, despite persistent geometric error.

By Zhenlin Yao, Wei Xiong
arXiv AI
Jul 7

Panorama: Fast-Track Nearest Neighbors

arXiv:2510. 00566v4 Announce Type: replace-cross Abstract: Approximate Nearest-Neighbor Search (ANNS) pipelines for high-dimensional neural embeddings spend the bulk of their query time in candidate verification, making it the primary bottleneck in the search process.

By Vansh Ramani, Alexis Schlomer, Akash Nayar, Sayan Ranu, Jignesh M. Patel, Panagiotis Karras
arXiv Machine Learning
Jun 4

On Out-of-sample Embedding in UMAP

arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.

By Mohammad Tariqul Islam, Jason W. Fleischer
arXiv Machine Learning
Aug 19

How smoothing the affinity matrix affects neighborhood preservation in t-SNE

The paper investigates how adjusting the sharpness of t‑SNE’s affinity matrix influences neighborhood preservation across scales. By applying a row‑wise power transform parameterized by γ, the authors can smooth or sharpen each row while keeping sparsity and rank order intact, effectively rescaling the Gaussian bandwidth and altering local perplexities. Experiments show that sharpening enhances the retention of the very nearest neighbors, whereas smoothing improves the preservation of broader local neighborhoods, outperforming existing multiscale affinity methods in the mid‑local range.

By Shirin Mohebi, Guillaume Bied, Jefrey Lijffijt