arXiv Machine Learning

On Out-of-sample Embedding in UMAP

arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.

arXiv Computer Vision
Sep 3

Aggregating Neighbor Embedding Projection and Rank-Based Manifold Learning for Image Retrieval

The paper introduces a new image retrieval framework that merges neighbor embedding projections with rank-based manifold learning via rank aggregation. It uses UMAP to create low‑dimensional feature representations and combines ranked lists from UMAP and rank‑based re‑ranking methods using the Borda Count strategy. Experiments on public datasets with ResNet152, Swin Transformer, and DINOv2 features show that this combined approach improves retrieval performance, especially in scenarios where baseline representations have low precision.

By Vinicius Atsushi Sato Kawai, Gustavo Rosseto Leticio, Lucas Pascotti Valem, Daniel Carlos Guimar\~aes Pedronette
arXiv Machine Learning
Aug 28

The Rashomon Effect for Visualizing High-Dimensional Data

The paper introduces the Rashomon set for dimension reduction, a collection of equally good embeddings that preserve high‑dimensional structure. It proposes PCA‑informed alignment to make axes interpretable, concept‑alignment regularization to incorporate external knowledge, and a method to extract trustworthy nearest‑neighbor relationships across the Rashomon set for refined embeddings. These techniques aim to produce interpretable, robust, and goal‑aligned visualizations by leveraging multiple valid embeddings instead of a single one.

By Yiyang Sun, Haiyang Huang, Gaurav Rajesh Parikh, Cynthia Rudin
arXiv Machine Learning
Sep 18

Out-of-Sample Embedding with Proximity Data: Projection versus Restricted Reconstruction

The paper reviews methods for adding a new point to a vector diagram using proximity data, a problem first examined by J.C. Gower in 1968. It classifies existing kernel-based approaches into two strategies: projection, analogous to adding a point in principal component analysis, and restricted reconstruction, which seeks to re‑optimize the multivariate analysis while keeping the existing diagram fixed. The authors show that each method can be derived from one of these two strategies and discuss when each strategy may be preferable.

By Michael W. Trosset, Kaiyi Tan, Minh Tang, Carey E. Priebe
arXiv Machine Learning
Sep 1

Effective Graph and Rank-based Contextual Embeddings for Textual and Multimedia Data

arXiv:2608.29001v1 Announce Type: new Abstract: In a data-driven world, efficiently organizing and mapping relationships between objects is crucial. Graphs are powerful tools for modeling these conne...

By Thiago C\'esar Castilho Almeida, Gustavo Rosseto Let\'icio, Lucas Pascotti Valem, Andr\'e Freitas, Daniel Carlos Guimar\~aes Pedronette