arXiv:2609.39816v1 Announce Type: new
Abstract: Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later t...
By Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie
arXiv:2608.31157v1 Announce Type: new
Abstract: Many parameter-efficient methods generate the parameters of a large neural network from a low-dimensional latent representation. Given an architecture...
By Shijun Zhang
arXiv:2511.02821v2 Announce Type: replace-cross
Abstract: We develop new accelerated first-order algorithms in the Frank-Wolfe (FW) family for minimizing smooth convex functions over compact convex s...
By Dan Garber
The paper proposes a new update geometry for the language‑model head by treating the head and softmax as a single module and using Hilbert’s projective distance to measure functional change. It replaces the spectral norm with the Euclidean row diameter, derives a tractable RowNorm update rule, and demonstrates that RowNorm substantially reduces step diameters and Hilbert perturbations while only slightly increasing validation loss.
By Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno, Zhengzhong Liu, Eric Xing
arXiv:2605.07815v2 Announce Type: replace
Abstract: Muon fixes the \emph{direction} of every matrix-valued update at the polar factor of its momentum, while each layer's step \emph{magnitude} is addr...
By Yuxuan Lou, Yang You
EGGROLL replaces dense Gaussian perturbations in evolution strategies with low‑rank Gaussian products, enabling practical optimization of large language models while maintaining exactness on quadratic objectives. The paper analyzes the mean update field, error bounds, and shows that rank‑one perturbations add only a small variance penalty compared to dense ES. A new leave‑one‑out estimator, LOO‑ROLL, further reduces estimator MSE and improves post‑training performance on transformer blocks and GSM8K accuracy.
By Ege C. Kaya, Abolfazl Hashemi
arXiv:2606. 20082v1 Announce Type: cross Abstract: The John ellipsoid of a symmetric polytope $P=\{\mathbf{x}\in\mathbb{R}^d:\|\mathbf{A}\mathbf{x}\|_\infty\le1\}$, $\mathbf{A}\in\mathbb{R}^{n\times d}$, is computed by a long line of leverage-score algorithms, from Cohen, Cousins, Lee and Yang (COLT 2019) to its successors [WY24, CLS+25], all reaching a $(1+\varepsilon)$-approximation in $\Theta(\varepsilon^{-1}\log(n/d))$ iterations.
By Xiaoyu Li, Junwei Yu, Jiaojiao Jiang, Junbin Gao, Andi Han
arXiv:2505. 21460v2 Announce Type: replace Abstract: We study online calibration of multi-dimensional forecasts over an arbitrary convex set $P \subset \mathbb{R}^d$ relative to an arbitrary norm $|\cdot|$.
By Maxwell Fishelson, Noah Golowich, Mehryar Mohri, Jon Schneider
EGGROLL makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractiv...
arXiv:2607. 05872v1 Announce Type: new Abstract: Memory-efficient optimizers such as GaLore train large language models by projecting gradients onto a rank-r subspace recomputed every T steps, assuming this subspace is a slowly drifting object that can be tracked.
By Noel Thomas
arXiv:2607. 16138v1 Announce Type: new Abstract: Improved Kernel Partial Least Squares (IKPLS) algorithms 1 and 2 are among the fastest PLS calibration algorithms.
By Ole-Christian Galbo Engstr{\o}m
arXiv:2603. 00910v2 Announce Type: replace-cross Abstract: Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas others are nearly redundant.
By Theophilus Amaefuna, Hitesh Vaidya, Anshuman Chhabra, Ankur Mali