arXiv Machine Learning

Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error

arXiv Machine Learning
Sep 11

Toward a First-Principles Update Geometry for the Language-Model Head

The paper proposes a new update geometry for the language‑model head by treating the head and softmax as a single module and using Hilbert’s projective distance to measure functional change. It replaces the spectral norm with the Euclidean row diameter, derives a tractable RowNorm update rule, and demonstrates that RowNorm substantially reduces step diameters and Hilbert perturbations while only slightly increasing validation loss.

By Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno, Zhengzhong Liu, Eric Xing
arXiv Machine Learning
Sep 11

EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale

EGGROLL replaces dense Gaussian perturbations in evolution strategies with low‑rank Gaussian products, enabling practical optimization of large language models while maintaining exactness on quadratic objectives. The paper analyzes the mean update field, error bounds, and shows that rank‑one perturbations add only a small variance penalty compared to dense ES. A new leave‑one‑out estimator, LOO‑ROLL, further reduces estimator MSE and improves post‑training performance on transformer blocks and GSM8K accuracy.

By Ege C. Kaya, Abolfazl Hashemi
arXiv Machine Learning
Jun 19

Beyond Averaging in John Ellipsoid Approximation: High-Accuracy Algorithms in the Leverage-Score Model

arXiv:2606. 20082v1 Announce Type: cross Abstract: The John ellipsoid of a symmetric polytope $P=\{\mathbf{x}\in\mathbb{R}^d:\|\mathbf{A}\mathbf{x}\|_\infty\le1\}$, $\mathbf{A}\in\mathbb{R}^{n\times d}$, is computed by a long line of leverage-score algorithms, from Cohen, Cousins, Lee and Yang (COLT 2019) to its successors [WY24, CLS+25], all reaching a $(1+\varepsilon)$-approximation in $\Theta(\varepsilon^{-1}\log(n/d))$ iterations.

By Xiaoyu Li, Junwei Yu, Jiaojiao Jiang, Junbin Gao, Andi Han
arXiv Machine Learning
Aug 12

High-Dimensional Calibration from Swap Regret

arXiv:2505. 21460v2 Announce Type: replace Abstract: We study online calibration of multi-dimensional forecasts over an arbitrary convex set $P \subset \mathbb{R}^d$ relative to an arbitrary norm $|\cdot|$.

By Maxwell Fishelson, Noah Golowich, Mehryar Mohri, Jon Schneider
arXiv Machine Learning
Jul 20

Improving Improved Kernel PLS

arXiv:2607. 16138v1 Announce Type: new Abstract: Improved Kernel Partial Least Squares (IKPLS) algorithms 1 and 2 are among the fastest PLS calibration algorithms.

By Ole-Christian Galbo Engstr{\o}m