arXiv:2606. 19878v1 Announce Type: new Abstract: Recent work on first-order optimizers for empirical risk minimization (ERM) has suggested that smoothness of ERM loss functions in the training data, rather than in the optimization parameters, can be leveraged to improve the oracle complexity of gradient descent (GD) methods.
By Dongmin Lee, William Lu, Anuran Makur
arXiv:2604.10857v2 Announce Type: replace-cross
Abstract: Diffusion models generate samples by iteratively querying learned score estimates. A rapidly growing literature focuses on accelerating sampl...
By Zhiyang Xun, Eric Price
arXiv:2604. 08625v2 Announce Type: replace-cross Abstract: We develop a theoretical framework for generalization in the interpolating regime of statistical learning.
By Gustav Olaf Yunus Laitinen-Lundstr\"om Fredriksson-Imanov
arXiv:2606. 01521v1 Announce Type: new Abstract: A central problem in machine learning is that models can achieve near-perfect training performance while generalizing substantially less well to unseen examples.
By Luca Muscarnera, Silas Ruhrberg Est\'evez, Yuanzhang Xiao, Mihaela Van der Schaar
arXiv:2109.02355v2 Announce Type: replace
Abstract: The last decade of progress in machine learning (ML), especially the deep learning era, has raised a number of scientific questions that challenge...
By Yehuda Dar, Vidya Muthukumar, Richard G. Baraniuk
arXiv:2606. 08554v1 Announce Type: new Abstract: This paper provides a theoretical account of memorization in stochastic interpolation models.
By Yunchen Li, Shaohui Lin, Zhou Yu
arXiv:2606. 07495v1 Announce Type: new Abstract: Understanding how training data shape neural network predictions is a central problem in modern learning theory.
By Jin Guo, Roy Y. He, Jean-Michel Morel
The paper develops a statistical theory for minimum‑norm interpolation in high‑dimensional regression, showing how regularization geometry and signal sparsity affect generalization. It identifies regimes where sparsity‑promoting regularizers yield exact interpolation that is far more accurate than approximate fitting, and proves a zero–one generalization law for strongly overparameterized noiseless problems. The authors also characterize training and generalization errors along ρ‑regularization paths when feature dimension and sample size are proportional, demonstrating that generalization improves with more sparsity‑promoting norms and sparser targets, and that small changes in regularization strength can cause large shifts in generalization.
whyItMatters":"The work provides a quantitative understanding of delayed generalization (grokking) and reveals a statistical instability in minimum‑norm interpolation, offering insights that could guide the design of regularizers for better generalization in overparameterized models."
By Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli
arXiv:2510. 06028v3 Announce Type: replace Abstract: This paper provides data-dependent bounds on the expected error of the Gibbs algorithm in the overparameterized interpolation regime, where low training errors are also obtained for impossible data, such as random labels in classification.
By Andreas Maurer, Erfan Mirzaei, Massimiliano Pontil
ChebBooster is a training‑free extrapolation framework that accelerates Diffusion Transformers (DiTs) by using Chebyshev polynomial theory. It employs a Barycentric formulation for numerically stable evaluation and separates the process into an offline weight precomputation phase and a lightweight online application stage. Experiments on DiT‑XL/2, PixArt‑Σ, and FLUX.1‑dev show consistent visual quality gains and up to 3.68× latency speedup and 5.12× FLOPs reduction compared to existing training‑free baselines.
By Chengjie Lu, Tianchi Deng, Zhengqi He, Chengwen Luo, Xueliang Li
arXiv:2509. 03758v5 Announce Type: replace Abstract: We propose a data-driven interpolation framework for reconstructing real-valued functions on smooth manifolds from scattered pointwise observations.
By Alvaro Almeida Gomez
arXiv:2606. 06469v1 Announce Type: cross Abstract: Let $S$ be the set of unit norm linear classifiers $\theta \in \mathbb{R}^d$ which correctly classify every point of a labeled dataset $(X_i,y_i)_{i=1}^n$, $X_i \in \mathbb{R}^d$, $y_i \in \{-1,+1\}$, with a possibly negative margin $\kappa$ fixed in advance.
By August Y. Chen, Ahmed El Alaoui