arXiv:2606. 05242v1 Announce Type: cross Abstract: Stochastic-gradient Langevin algorithms often use tamed denominators to stabilize non-globally Lipschitz drifts.
By Yiwei Zhou, Ziheng Chen
arXiv:2607. 19544v1 Announce Type: cross Abstract: We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift.
By Yiwei Zhou, Ziheng Chen
We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift. A threshold determines where the taming turns on, while a relative-growth principle derived from the one-step Lyapunov stability condition determines the required taming strength.
arXiv:2607. 08104v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) is a cornerstone of modern optimization.
By Ryusei Yamada, Naoki Sato, Hideaki Iiduka
arXiv:2609.06064v1 Announce Type: cross
Abstract: Stochastic min-max optimization has attracted increasing attention due to its applications in modern machine learning, while existing theoretical stu...
By Tianxi Zhu, Yi Xu, Xiangyang Ji
The paper introduces SHANG++—an accelerated stochastic gradient descent algorithm designed to be robust under multiplicative noise scaling (MNS). Building on a semi‑implicit discretization called SHANG, SHANG++ adds a damping correction that improves stability and convergence for both convex and strongly convex objectives. Experiments on convex problems and deep learning tasks, including a noise‑robust test on ResNet‑34, show that SHANG++ consistently outperforms existing accelerated methods with minimal parameter sensitivity.
By Yaxin Yu, Long Chen, Minfu Feng
arXiv:2601. 12238v5 Announce Type: replace-cross Abstract: In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness.
By Sharan Sahu, Cameron J. Hogan, Martin T. Wells
The paper proves that stochastic gradient descent with gradient clipping and additive Gaussian noise (SGD‑CN) converges almost surely under smoothness and bounded noise assumptions, given standard decaying step sizes. The analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, showing that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods in both convex and nonconvex regimes.
By Amartya Mukherjee, Jun Liu
arXiv:2505.20817v3 Announce Type: replace-cross
Abstract: Gradient clipping is widely used in language-model training to control heavy-tailed gradient noise and can improve convergence guarantees ove...
By Taha El Bakkali El Kadi, Savelii Chezhegov, Aleksandr Beznosikov, Samuel Horv\'ath, Eduard Gorbunov
arXiv:2607. 15412v1 Announce Type: new Abstract: Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously.
By Chentong Huang, Lisha Chen
The paper introduces penalized nonreversible Langevin algorithms for sampling from a target distribution constrained to a compact convex set. It combines a squared distance penalty with skew-symmetric perturbations that preserve the penalized Gibbs distribution, and provides nonasymptotic total variation and Wasserstein bounds under various smoothness and contraction assumptions. Numerical experiments demonstrate the methods on constrained Bayesian regression, classification, neural networks, and truncated sampling, highlighting acceleration in a stochastic quadratic model.
By Pervez Ali, Weihao Dong, Xiaoyu Wang
The paper presents a theoretical study of Adam in non‑stationary stochastic optimization, distinguishing two regimes: Euclidean tracking under adaptive strong monotonicity and high‑probability projected stationarity for general smooth objectives. It derives finite‑time bounds that decompose into initialization, objective drift, first‑moment tracking error (β₁), and preconditioner perturbation (β₂), and characterizes burn‑in times for constant and step‑decay schedules. The analysis reveals a noise–drift tradeoff, showing that in noise‑dominated settings Adam’s adaptive mechanisms can improve guarantees, while in drift‑dominated settings they may worsen tracking, potentially making vanilla SGD preferable.
By Sharan Sahu, Abir Sarkar, Cameron J. Hogan, Martin T. Wells