arXiv:2609.39512v1 Announce Type: new
Abstract: The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation a...
By Hong Zheng
arXiv:2605.28517v2 Announce Type: replace-cross
Abstract: Stochastic gradient descent with momentum (SGDM) is one of the most widely used optimization algorithms in machine learning. While optimizati...
By Yunwen Lei, Zimeng Wang, Xiaoming Yuan
Uniform stability controls how much one training example can change the loss at any test point. A new logarithmic-free upper bound shows that a $γ$-uniformly stable algorithm with loss in $[0,L]$ has...
arXiv:2604. 10727v2 Announce Type: replace-cross Abstract: Classical information-theoretic learning bounds typically rely on KL mutual information and moment-generating-function (MGF) arguments, which are well matched to bounded or sub-Gaussian losses but can be ineffective when losses or rewards are heavy-tailed.
By Huiming Zhang, Binghan Li, Wan Tian, Qiang Sun
arXiv:2601.11701v2 Announce Type: replace-cross
Abstract: Algorithmic stability is a central concept in statistics and learning theory that measures how sensitive an algorithm's output is to small ch...
By Abhinav Chakraborty, Yuetian Luo, Rina Foygel Barber
The paper discusses the data processing inequality (DPI) in statistics, which states that a stochastically modified experiment cannot have a lower Bayes risk than the original. It shows that this classical DPI does not hold for constrained learning problems common in machine learning, where the model class is limited. The authors propose a generalized DPI that applies to constrained Bayes risks, linking it to a set containment condition on a superprediction set, and provide sufficient conditions for this containment.
By Laura Iacovissi, Rabanus Derr, Robert C. Williamson
arXiv:2602. 05657v2 Announce Type: replace Abstract: The study of tail behaviour of SGD-induced processes has been attracting a lot of interest, due to offering strong guarantees with respect to individual runs of an algorithm.
By Aleksandar Armacki, Dragana Bajovi\'c, Du\v{s}an Jakoveti\'c, Soummya Kar, Ali H. Sayed
The paper presents a polynomial‑time algorithm for robustly learning Boolean concept classes with respect to a fixed distribution, achieving the optimal error rate of η + ε where η is the noise rate. It builds on Blanc’s earlier, computationally inefficient algorithm and introduces no‑regret learners to overcome the previous limitations. Additionally, the authors provide an efficient method that does not require an ERM oracle for any function class admitting sandwiching polynomials under hypercontractive distributions, including a first polynomial‑time solution for learning halfspaces with Gaussian marginals at error η + ε.
By Adam R. Klivans, Konstantinos Stavropoulos, Sergei Tikhonov, Arsen Vasilyan
arXiv:2602. 20971v3 Announce Type: replace-cross Abstract: Bubeck and Selke (2021) propose the connection between the Law of Robustness and robust generalization error as an open problem.
By Mihir More, Aritra Das, Jaee Ponde, Himadri Mandal, Vishnu Varadarajan, Debayan Gupta
arXiv:2608.24098v1 Announce Type: new
Abstract: Uniform stability controls how much one training example can change the loss at any test point. A new logarithmic-free upper bound shows that a $\gamma...
By Pahan Dewasurendra
The paper establishes a shrinking‑tube concentration bound for projected stochastic approximation driven by an adaptive Markov chain, guaranteeing that after a chosen time every iterate stays within a tolerance that tightens over time. The bound’s probability of any exit after that time decays polynomially, and a matching lower bound shows this exponent is optimal under finite second moments. Extensions to recursions with martingale‑difference noise and predictable bias reveal how noise scale and bias affect exit‑probability decay and tube shrinkage, with applications to inventory learning and numerical gradient accuracy.
By Jin Li, Ye Luo, Xiaowei Zhang
arXiv:2606. 00520v1 Announce Type: cross Abstract: Many stochastic gradient methods are believed not to converge when the noise in stochastic gradients has only a finite $p$-th moment for $p\in\left(1,2\right)$, a setting known as the heavy-tailed noise assumption.
By Zijian Liu