arXiv Machine Learning By Louise Davy, Stephan Cl\'emen\c{c}on, Charlotte Laclau

Doing well with less! On Sampling Techniques for Empirical Pairwise Loss Estimation/Minimization

Read the original on arXiv Machine Learning →

arXiv:2606. 02345v1 Announce Type: cross Abstract: Many machine learning problems, including similarity learning, ranking, and clustering, rely on empirical pairwise loss functions whose quadratic computational cost quickly becomes prohibitive at scale.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 7

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

arXiv:2606. 12913v2 Announce Type: replace Abstract: The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning~(DP) methods which retain only a subset of informative samples to reduce training cost.

By Dongyue Wu, Zilin Guo, Xiaoyu Li, Jiajia Liu, Jingdong Chen, Nong Sang, Changxin Gao
arXiv Machine Learning
Jun 4

On Out-of-sample Embedding in UMAP

arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.

By Mohammad Tariqul Islam, Jason W. Fleischer
arXiv AI
Sep 24

Scalable Subgraph Sampling via Resistance Curvature

The paper introduces a scalable subgraph sampling method that uses resistance curvature to guide the selection of nodes and edges for graph neural network training. It builds on ERC‑LG, a curvature approximation technique that employs Johnson‑Lindenstrauss projections and regularized multi‑GPU batched conjugate gradient solvers, thereby avoiding costly Laplacian pseudoinverse calculations and large embedding storage. Experiments demonstrate that ERC‑LG‑based sampling matches pseudoinverse‑based curvature numerically, runs faster than conjugate‑gradient‑only approaches, and achieves the best mean accuracy on six of seven real‑world node‑classification datasets.

By Chaoqun Fei, Tinglve Zhou, Tianyong Hao, Yangyang Li
arXiv AI
Sep 10

Revisiting Thinning Methods for Kernel Learning Problems

The paper introduces Backward Kernel Herding, an algorithm that iteratively removes data points to create representative subsets for kernel learning, achieving performance comparable to state‑of‑the‑art methods while speeding up subsampling when the reduced size is less than half the original dataset. It also proposes Flexible Kernel Thinning, an extension that allows construction of subsets of any size, not just successive halvings, and demonstrates that this method often yields the best predictive performance. Experiments on Gaussian Processes and Kernel Support Vector Machines show that Backward Kernel Herding excels in training‑time efficiency, while Flexible Kernel Thinning offers superior predictive accuracy and competitive memory usage, emphasizing the need to choose a reduction strategy based on the desired trade‑off between performance, cost, and memory.

By Blanca Cano-Camarero, Yago R. Aguado-Carrillo-de-Albornoz, \'Angela Fern\'andez-Pascual, Jos\'e R. Dorronsoro