On the Principles Behind Neural Network Optimizers
arXiv:2608. 16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation.
The paper introduces a novel generalization of the Adam optimizer to manifold settings, specifically targeting homogeneous spaces such as the Stiefel, symplectic Stiefel, and Grassmann manifolds. By exploiting a global tangent space representation (the Lie subspace), the authors eliminate the need for projection steps and enable all Adam operations to be performed directly on these manifolds. The new optimizer is applied to train transformers and a symplectic autoencoder, achieving orthogonality constraints to machine precision and outperforming existing methods.
arXiv:2608. 16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation.
arXiv:2609.21039v1 Announce Type: new Abstract: A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form $W = BA$ in which two parameter matri...
arXiv:2606. 04623v2 Announce Type: replace Abstract: High-dimensional Hamiltonian systems play a central role in many scientific and engineering disciplines, with dynamics that evolve on symplectic manifolds.
arXiv:2606. 04623v1 Announce Type: new Abstract: High-dimensional Hamiltonian systems play a central role in many scientific and engineering disciplines, with dynamics evolving on symplectic manifolds.
arXiv:2609.35436v2 Announce Type: replace Abstract: Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications....
arXiv:2608.29867v1 Announce Type: new Abstract: Autoencoders are widely used for nonlinear dimensionality reduction and manifold learning. While most common implementations rely on both nonlinear enc...
arXiv:2604. 20308v2 Announce Type: replace Abstract: Graph neural networks face two fundamental challenges rooted in the linear structure of Euclidean vector spaces: (1) Current architectures represent geometry through vectors (directions, gradients), yet many tasks require matrix-valued representations that capture relationships between directions-such as how atomic orientations covary in a molecule.
arXiv:2605. 28983v2 Announce Type: replace-cross Abstract: In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights.
arXiv:2607. 19305v1 Announce Type: cross Abstract: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations.
arXiv:2602. 22895v2 Announce Type: replace-cross Abstract: Implementations of symmetric positive definite (SPD) matrix-based neural networks for neural decoding remain fragmented across research codebases and Python packages.
arXiv:2607. 08783v1 Announce Type: cross Abstract: Manifold-valued measurements are prevalent in various machine learning tasks.
arXiv:2607. 19305v2 Announce Type: replace-cross Abstract: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations.