arXiv:2609. 20687v1 Announce Type: cross Abstract: We study first-order black-box convex optimization over an $\ell_p$-ball for objectives Lipschitz in the $\ell_q$-norm, solving in the affirmative the nonsmooth version of the COLT open question (Guz15b) on whether the geometry of a smaller feasible set ($p < q$) can improve convergence rates in convex optimization, and matching prior lower bounds up to logarithmic factors.
By David Mart\'inez-Rubio, Brian Bullins, Crist\'obal Guzm\'an, Mathieu Molina
This paper investigates the ρ^p-Lipschitz constants of deep ReLU neural networks with random weights drawn from a He‑style initialization. For zero‑bias networks, it provides high‑probability upper and lower bounds that differ by at most a logarithmic factor in depth, and shows a sharp contrast between the regimes p∈[1,2) and p∈[2,∞], with the former behaving like the Euclidean norm of a Gaussian vector and the latter like its dual norm. The analysis is extended to networks with non‑zero biases from symmetric distributions, yielding bounds that differ by a logarithmic factor in width and a linear factor in depth.
By Sjoerd Dirksen, Patrick Finke, Paul Geuchen, Dominik St\"oger, Felix Voigtlaender
arXiv:2406. 13944v2 Announce Type: replace-cross Abstract: This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available.
By Yanke Song, Kenneth Gu, Sohom Bhattacharya, Pragya Sur
arXiv:2609. 03129v1 Announce Type: cross Abstract: Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks.
By Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
arXiv:2511. 11498v2 Announce Type: replace-cross Abstract: We consider the problems of \emph{learning} and \emph{testing} real-valued convex functions over Gaussian space.
By Renato Ferreira Pinto Jr., Cassandra Marcussen, Elchanan Mossel, Shivam Nadimpalli
The paper develops a statistical theory for minimum‑norm interpolation in high‑dimensional regression, showing how regularization geometry and signal sparsity affect generalization. It identifies regimes where sparsity‑promoting regularizers yield exact interpolation that is far more accurate than approximate fitting, and proves a zero–one generalization law for strongly overparameterized noiseless problems. The authors also characterize training and generalization errors along ρ‑regularization paths when feature dimension and sample size are proportional, demonstrating that generalization improves with more sparsity‑promoting norms and sparser targets, and that small changes in regularization strength can cause large shifts in generalization.
whyItMatters":"The work provides a quantitative understanding of delayed generalization (grokking) and reveals a statistical instability in minimum‑norm interpolation, offering insights that could guide the design of regularizers for better generalization in overparameterized models."
By Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli