arXiv:2407. 01718v2 Announce Type: replace-cross Abstract: Embedding high-dimensional data into a low-dimensional space is an indispensable component of data analysis.
By Boris Landa, Yuval Kluger, Rong Ma
arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.
By Mohammad Tariqul Islam, Jason W. Fleischer
arXiv:2601. 20844v3 Announce Type: replace-cross Abstract: This paper studies the Minimal Embeddable Dimension (MED): the least dimension in which there exists a configuration of $m$ object vectors so that every subset of size at most $k$ is exactly retrieved by score comparison.
By Zihao Wang, Hang Yin, Lihui Liu, Hanghang Tong, Yangqiu Song, Ginny Wong, Simon See
The paper introduces the Sparse Landmark Embedding (SLE) kernel, a new framework that removes the need for conditionally negative definite (CND) distance measures in kernel methods and Gaussian Processes. By embedding each input into a sparse feature vector using compactly supported bump functions centered at all training points, any standard positive semi-definite (PSD) kernel can be applied in this embedding space, guaranteeing PSD for arbitrary distance measures. The authors provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and show through experiments with geodesic and Wasserstein distances that the SLE kernel matches or surpasses domain-specific baselines in predictive accuracy and uncertainty quantification.
By Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
arXiv:2606. 16045v1 Announce Type: new Abstract: In the data selection problem, the objective is to choose a small, representative subset of data that can be used to efficiently train a machine learning model.
By Vincent Cohen-Addad, Sasidhar Kunapuli, Vahab Mirrokni, Mahdi Nikdan, David P. Woodruff, Samson Zhou
arXiv:2609. 02155v1 Announce Type: new Abstract: The Johnson-Lindenstrauss (JL) lemma guarantees that a random projection of $n$ points to $m=O(\varepsilon^{-2}\log n)$ dimensions preserves pairwise squared distances within relative error $\varepsilon$ with high probability, and this dimension order is asymptotically optimal.
By Piyush Sao