arXiv:2609.38095v1 Announce Type: new
Abstract: Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU...
By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e
The paper introduces an online sketched Newton method that uses a generalized accelerated sketch-and-project solver (GAS) to approximate Newton directions efficiently. GAS incorporates Nesterov momentum and a flexible projection metric, achieving accelerated convergence and reduced computational cost. The authors prove asymptotic normality and a functional central limit theorem for the averaged iterates, enabling an online inference procedure via random scaling that yields a pivotal test statistic with a parameter‑free limiting distribution.
By Xinchen Du, Elizaveta Rebrova, Micha{\l} Derezi\'{n}ski, Sen Na
arXiv:2609.36692v1 Announce Type: cross
Abstract: Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram represen...
By Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu
arXiv:2606. 01521v1 Announce Type: new Abstract: A central problem in machine learning is that models can achieve near-perfect training performance while generalizing substantially less well to unseen examples.
By Luca Muscarnera, Silas Ruhrberg Est\'evez, Yuanzhang Xiao, Mihaela Van der Schaar
arXiv:2402.11215v4 Announce Type: replace
Abstract: The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-sc...
By Tim Tsz-Kit Lau, Han Liu, Mladen Kolar
arXiv:2511. 19716v3 Announce Type: replace-cross Abstract: Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise.
By Mitchell Scott, Tianshi Xu, Ziyuan Tang, Alexandra Pichette-Emmons, Qiang Ye, Yousef Saad, Yuanzhe Xi