arXiv:2608. 26052v1 Announce Type: new Abstract: Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task.
By Gerard Conangla Planes
arXiv:2607. 23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it?
By Byeong Hoon Yoon
arXiv:2602. 18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation.
By Seyed Morteza Emadi
arXiv:2609.06327v2 Announce Type: replace-cross
Abstract: A query-oblivious coreset for a softmax-attention head is a subset of the key-value pairs whose attention output is within $\varepsilon$ of t...
By Ofek I. Cohen
The paper proposes a new update geometry for the language‑model head by treating the head and softmax as a single module and using Hilbert’s projective distance to measure functional change. It replaces the spectral norm with the Euclidean row diameter, derives a tractable RowNorm update rule, and demonstrates that RowNorm substantially reduces step diameters and Hilbert perturbations while only slightly increasing validation loss.
By Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno, Zhengzhong Liu, Eric Xing
arXiv:2609.09130v1 Announce Type: new
Abstract: An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input...
By Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang, Junbin Gao, Andi Han