arXiv:2609.23163v1 Announce Type: cross
Abstract: Comparing probability measures in machine learning trades transport geometry against computational cost: Wasserstein distances encode the geometry of...
By Mehrdad Mohammadi
arXiv:2609. 17845v1 Announce Type: cross Abstract: Let the exact homogeneous hard-margin support vector machine be trained on \(m\) independent observations from a Borel probability law on a real Hilbert space.
By Steve Hanneke, Aryeh Kontorovich
Let the exact homogeneous hard-margin support vector machine be trained on \(m\) independent observations from a Borel probability law on a real Hilbert space. We prove that, with score zero counted a...
arXiv:2607. 03815v1 Announce Type: cross Abstract: For compact convex sets $L,K \subset \mathbb{R}^n$, denote by $\lambda_K(L)$ the smallest size of a homothet of $K$ that contains $L$.
By Egor Bakaev, Amir Yehudayoff
arXiv:2609.38834v1 Announce Type: cross
Abstract: Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negativ...
By Dionysis Arvanitakis, Vaggos Chatziafratis, Yiyuan Luo, Konstantin Makarychev
arXiv:2609.06327v2 Announce Type: replace-cross
Abstract: A query-oblivious coreset for a softmax-attention head is a subset of the key-value pairs whose attention output is within $\varepsilon$ of t...
By Ofek I. Cohen
arXiv:2602. 19172v2 Announce Type: replace Abstract: Realizable online regression can behave very differently from online classification.
By Ilan Doron-Arad, Idan Mehalel, Elchanan Mossel
arXiv:2609. 20687v1 Announce Type: cross Abstract: We study first-order black-box convex optimization over an $\ell_p$-ball for objectives Lipschitz in the $\ell_q$-norm, solving in the affirmative the nonsmooth version of the COLT open question (Guz15b) on whether the geometry of a smaller feasible set ($p < q$) can improve convergence rates in convex optimization, and matching prior lower bounds up to logarithmic factors.
By David Mart\'inez-Rubio, Brian Bullins, Crist\'obal Guzm\'an, Mathieu Molina
An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al.
arXiv:2609.09130v1 Announce Type: new
Abstract: An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input...
By Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang, Junbin Gao, Andi Han
arXiv:2608.30254v1 Announce Type: new
Abstract: We resolve the threshold part of Question 4 of the COLT 2025 open problem "Data Selection for Regression Tasks" of Hanneke, Moran, Shlimovich and Yehud...
By Guangjian Zhang
arXiv:2609.15179v1 Announce Type: cross
Abstract: The Gaussian kernel is a widely used similarity measure underlying kernel methods such as kernel PCA and spectral clustering, but computing Gaussian...
By Soumik Dutta, Kunal Dutta