arXiv Machine Learning

EMS Coreset: An Efficient Expectation-Maximization Algorithm for Sinkhorn Coreset

arXiv:2608. 16101v1 Announce Type: cross Abstract: Coresets distill large datasets into small, representative subsets for efficient downstream learning.

arXiv Machine Learning
Jun 25

Sample complexity of unbalanced entropic OT

arXiv:2606. 24987v1 Announce Type: cross Abstract: Optimal transport (OT) has become a central language for comparing probability measures, but exact balanced OT is often both too rigid for data with missing, created, or destroyed mass and subject to unfavorable high-dimensional sample complexity.

By Francisco Andrade, Gabriel Peyr\'e, Clarice Poon
arXiv Machine Learning
1d ago

Low-Budget Active Learning through Entropic Optimal Transport

The paper introduces a low-budget active learning approach that selects a small coreset of data points for training high-accuracy models, particularly useful when labeling is expensive, such as in medical contexts. It uses features from a pretrained self-supervised model and applies entropic optimal transport—specifically the Sinkhorn divergence—as the selection criterion, enabling dimension-free sample complexity and efficient gradient-based optimization. The method combines gradient-based candidate generation with a swap-based local search, achieving superior performance over existing heuristics on image and medical datasets.

By Rim Hajal, Mathieu Besan\c{c}on, J\'er\^ome Malick
arXiv AI
Sep 10

Deep Barycentric Regression for Optimal Transport Map Estimation and its Statistical Optimality

The paper introduces BROT, a two‑step approach for estimating optimal transport maps. First, it computes the unregularized OT plan, then fits a deep neural network to the resulting barycentric targets using least‑squares regression. The authors prove that, under standard regularity conditions, BROT achieves the minimax convergence rate when the true OT map is Lipschitz, and demonstrate its effectiveness on synthetic data, images, and downstream tasks such as single‑cell perturbation prediction and unsupervised domain adaptation.

By Kunwoong Kim, Insung Kong, Yongdai Kim