arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
arXiv:2606. 30226v1 Announce Type: new Abstract: Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimization dynamics.
By Marcelina Marjankowska, Valerio Modugno, Paolo Barucca
arXiv:2607. 09967v1 Announce Type: cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps.
By Ethan Smith
arXiv:2606. 10929v1 Announce Type: cross Abstract: Task vectors, LoRA, activation steering, and random search around pretrained weights all suggest that learned behaviour can be controlled by linear directions.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2510.17072v2 Announce Type: replace
Abstract: Regression with non-Euclidean responses---e.g., probability distributions, networks, symmetric positive-definite matrices, and compositions---has b...
By Kyum Kim, Yaqing Chen, Paromita Dubey
arXiv:2609.38081v1 Announce Type: new
Abstract: On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these...
By Ann Huang, Mitchell Ostrow, Zhouyang Lu, William T. Redman, Leo Kozachkov, Kanaka Rajan