arXiv:2606. 08638v1 Announce Type: cross Abstract: Recent research has developed practical, parallelizable first-order methods for large scale linear programming, but performance is highly dependent on hyperparameter selection.
By Siddharth Prasad, Dravyansh Sharma
arXiv:2602. 14656v2 Announce Type: replace Abstract: Orthogonality constraints are ubiquitous in robust and probabilistic machine learning.
By Adri\'an Javaloy, Antonio Vergari
The paper introduces a GPU-resident, batched Levenberg–Marquardt solver that efficiently optimizes constants in tree-based genetic programming for symbolic regression. By using reverse-mode automatic differentiation to assemble per-tree Jacobians in a single backward sweep, the solver’s per-iteration cost becomes independent of the number of constants per tree, achieving up to 510,000 trees per second on an NVIDIA A100. Integrated into EvoGP, the solver enables end-to-end search that recovers governing equations on 10 of 18 constructed problems, a significant improvement over stock EvoGP.
By Hao Mao, Xu Tony Liu, Shuai Lu, Peng Zhao, Wenzheng Jiang, Yuntian Chen
arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.
By Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang
arXiv:2609.38095v1 Announce Type: new
Abstract: Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU...
By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e
arXiv:2602. 02016v2 Announce Type: replace Abstract: Shampoo is one of the leading approximate second-order optimizers: a variant of it has won the MLCommons AlgoPerf competition, and it has been shown to produce models with lower activation outliers that are easier to compress.
By Ionut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan, Dan Alistarh
arXiv:2606. 07574v1 Announce Type: cross Abstract: Manifold-constrained hyper-connections (mHCs) have recently been proposed as a principled extension of hyper-connections, where the residual mixing matrices are constrained to be doubly stochastic via projection onto the Birkhoff polytope.
By Chenrui Wang, Yixuan Qiu
arXiv:2512. 02551v3 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels.
By Songqiao Su, Xiaoya Li, Albert Wang, Guoyin Wang, Jiwei Li, Chris Shum
arXiv:2606. 18463v1 Announce Type: cross Abstract: Distributed stochastic gradient descent (SGD) is limited by communication rather than computation, since each iteration requires an AllReduce across processes.
By Aditya Devarakonda, Irene Sim\'o Mu\~noz, Giulia Guidi
arXiv:2601. 13994v3 Announce Type: replace-cross Abstract: Differentiable sparse linear algebra is foundational for scientific machine learning, yet PyTorch lacks a unified library for it: torch.
By Mingyuan Chi, Shizheng Wen
arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.
By Nadav Benedek, Tomer Koren, Ohad Fried
arXiv:2609.36692v1 Announce Type: cross
Abstract: Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram represen...
By Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu