arXiv Machine Learning

The kernel of graph indices for vector search

The paper introduces the Support Vector Graph (SVG), a graph index for vector search that uses kernel methods to guarantee navigability in both metric and non‑metric vector spaces, such as inner product similarity. It shows that popular indices like HNSW and DiskANN are special cases of SVG, and proposes SVG‑L0, which adds an ℓ₀ sparsity constraint to enforce bounded out‑degree while maintaining computational efficiency.

arXiv Machine Learning
Jun 10

$k$-Nearest Neighbors in Gromov--Wasserstein Space

arXiv:2606. 10295v1 Announce Type: cross Abstract: The Gromov--Wasserstein (GW) distance provides a framework for comparing metric measure spaces, regardless of their underlying structure or geometry.

By Kaitlyn Hohmeier, Nicolas Fraiman, Caroline Moosmueller
arXiv Machine Learning
Jun 18

Compact Geometric Representations of Hierarchies

arXiv:2606. 18520v1 Announce Type: cross Abstract: Computing geometric representations of data is a cornerstone of modern machine learning, typically achieved by training dual encoders which map queries and documents into a shared embedding space.

By Prashant Gokhale, Piotr Indyk, Yuhao Liu, Sandeep Silwal, Tony Chang Wang, Haike Xu
arXiv Machine Learning
Sep 4

Geometry-Aware Graph Construction via Adaptive Spectral Bandwidth Control

The paper introduces a geometry‑aware graph construction method that adaptively selects Gaussian kernel bandwidths per node to align the kernel’s spectral complexity with the intrinsic dimensionality of the underlying manifold. By matching the kernel’s effective rank to a local intrinsic dimension estimate derived from a minimum spanning tree, the method operates within a manifold‑consistent log‑log scaling regime. Experiments on CIFAR‑100 demonstrate that this adaptive bandwidth approach consistently improves leave‑one‑out classification and label propagation accuracy compared to fixed‑bandwidth and other adaptive techniques.

By Ecem Bozkurt, Antonio Ortega
arXiv Machine Learning
Aug 27

Optimal Time Complexity Algorithms for Computing General Random Walk Graph Kernels on Sparse Graphs

The paper introduces linear‑time randomized algorithms for unbiased approximation of general random walk kernels (RWKs) on sparse graphs, covering both labelled and unlabelled cases. By sampling dependent random walks and constructing novel graph embeddings in ρ^d, the method avoids building the direct product graph, enabling scaling to massive datasets that cannot fit on a single machine. The authors provide exponential concentration bounds for the estimator’s sharpness and demonstrate up to 27× speed‑ups and 128× larger graph handling compared to previous cubic‑time approaches.

By Krzysztof Choromanski, Isaac Reid, Arijit Sehanobish, Avinava Dubey
arXiv Machine Learning
Jul 7

HNSW with Accuracy Guarantees Using Graph Spanners

arXiv:2607. 02338v2 Announce Type: replace-cross Abstract: Hierarchical Navigable Small World (HNSW) graphs serve as the industry standard due to their logarithmic complexity and strong empirical performance.

By Minghao Li, Raghav Mittal, Sanjivni Rana, Suraj Shetiya, Gautam Das, Nick Koudas