arXiv Machine Learning

Inexact calculus of variations on the hyperspherical tangent bundle with connections to the attention mechanism

arXiv:2507. 15431v4 Announce Type: replace Abstract: We offer a theoretical mathematical background through Lagrangian optimization on the unit hyperspherical manifold and its tangential structure.

arXiv AI
Aug 5

Sphere Retraction Normalizations

arXiv:2608. 02668v1 Announce Type: cross Abstract: Residual connections are the de facto mechanism for training deep neural networks stably.

By Jie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
arXiv Machine Learning
Aug 20

The Diffusion-Attention Connection

arXiv:2604. 09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain.

By Julio Candanedo
arXiv Machine Learning
Sep 17

Generalizing Adam to Manifolds for Efficiently Training Transformers

The paper introduces a novel generalization of the Adam optimizer to manifold settings, specifically targeting homogeneous spaces such as the Stiefel, symplectic Stiefel, and Grassmann manifolds. By exploiting a global tangent space representation (the Lie subspace), the authors eliminate the need for projection steps and enable all Adam operations to be performed directly on these manifolds. The new optimizer is applied to train transformers and a symplectic autoencoder, achieving orthogonality constraints to machine precision and outperforming existing methods.

By Benedikt Brantner
arXiv Machine Learning
Aug 18

Iso-Riemannian Optimization on Learned Data Manifolds

arXiv:2510. 21033v3 Announce Type: replace-cross Abstract: We develop a theory of iso-Riemannian optimization for problems constrained to learned data manifolds, a setting in which classical Riemannian optimization - and Riemannian gradient descent in particular - can be poorly suited.

By Willem Diepeveen, Melanie Weber
arXiv Machine Learning
Sep 21

Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs

The paper extends the analysis of Joint-Embedding Predictive Architectures (JEPAs) beyond Euclidean latent spaces to Riemannian manifolds. It shows that when latent variables lie on a sphere and the target distribution matches this spherical geometry, every optimal representation recovers the latent state up to an orthogonal transformation, demonstrating that Gaussian uniqueness is not universal. Experiments confirm that geometrically compatible targets improve linear recovery, especially in high-dimensional toroidal settings.

By L\'eo Nicollier (CB, ATT), Enric Meinhardt-Llopis (CB), Marc Pic (ATT), Pablo Mus\'e (CB, IFUMI), Gabriele Facciolo (CB)
arXiv Machine Learning
Sep 24

hyperbolix: Hyperbolic Deep Learning in JAX

hyperbolix is an open‑source library for hyperbolic deep learning in JAX, built on Flax NNX. It provides six manifolds—including Euclidean, Poincaré ball, hyperboloid, κ‑stereographic, mixed‑curvature product, and proper velocity space—through a common interface, and implements a wide range of layer families (linear, convolution, attention, normalization, positional encoding, regression, vector quantization). The library also supplies Riemannian optimizers, wrapped distributions, dimensionality‑reduction techniques, and precision‑tested operations that replace numerically unstable formulas on the hyperboloid, ensuring accurate float32 computations at large distances.

By Timo Klein, Thomas Lang, Yllka Velaj, Sebastian Tschiatschek