arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
arXiv:2606. 30226v1 Announce Type: new Abstract: Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimization dynamics.
By Marcelina Marjankowska, Valerio Modugno, Paolo Barucca
arXiv:2607. 09967v1 Announce Type: cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps.
By Ethan Smith
arXiv:2606. 10929v1 Announce Type: cross Abstract: Task vectors, LoRA, activation steering, and random search around pretrained weights all suggest that learned behaviour can be controlled by linear directions.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2510.17072v2 Announce Type: replace
Abstract: Regression with non-Euclidean responses---e.g., probability distributions, networks, symmetric positive-definite matrices, and compositions---has b...
By Kyum Kim, Yaqing Chen, Paromita Dubey
arXiv:2609.38081v1 Announce Type: new
Abstract: On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these...
By Ann Huang, Mitchell Ostrow, Zhouyang Lu, William T. Redman, Leo Kozachkov, Kanaka Rajan
arXiv:2608.24568v1 Announce Type: cross
Abstract: Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with the geomet...
By Paul Caillon, Christophe Cerisara, Alexandre Allauzen
arXiv:2606. 00442v1 Announce Type: new Abstract: Many machine learning techniques rely on approximating a loss function's curvature, but this is notoriously hard to do at the scale of modern deep networks.
By Artem Artemev, Rui Xia, Benjamin M. Boyd, Youjing Yu, Felix Dangel, Guillaume Hennequin, Alberto Bernacchia
AYLA is a loss reparameterization framework that applies a sigmoid‑controlled power‑law transformation to the empirical loss, dynamically adjusting gradient magnitudes without changing stationary points or optimal solutions. By reshaping optimization trajectories, AYLA accelerates descent in flat or saddle‑dominated regions and stabilizes late‑stage training, leading to improved feature recovery in two‑layer tanh networks on synthetic Gaussian data. Experiments show enhanced weight alignment, neuron similarity, activation correlation, and richer internal representations, while mitigating rank collapse and promoting a transition from lazy to active feature‑learning regimes.
By Behnam Gheshlaghi, Shahin Atakishiyev
arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.
By Tuc Nguyen, Thai Le
arXiv:2603. 29824v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning methods such as LoRA enable efficient adaptation of large pretrained models, but often lag behind full fine-tuning in both convergence speed and final performance.
By Fr\'ed\'eric Zheng, Alexandre Prouti\`ere
The paper introduces a new way to evaluate materials graph neural networks (GNNs) by measuring how many trainable parameter‑space directions are needed to achieve good performance. Using random‑subspace intrinsic‑dimension analysis, the authors train CGCNN, ALIGNN, and DimeNet++ on six prediction tasks and plot recovery curves that separate final accuracy from the dimensional demand required to reach it. The study finds that different tasks and architectures vary in how sensitive they are to dimensional restriction, revealing insights that final error metrics alone miss.
By Shehroz Ahmad Shoaib, Kangming Li