arXiv:2608. 02668v1 Announce Type: cross Abstract: Residual connections are the de facto mechanism for training deep neural networks stably.
By Jie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
arXiv:2604. 09560v2 Announce Type: replace Abstract: Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain.
By Julio Candanedo
arXiv:2608. 10416v1 Announce Type: cross Abstract: We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realization (Riemann GeoResolver).
By Liangchen Ge
The paper introduces a novel generalization of the Adam optimizer to manifold settings, specifically targeting homogeneous spaces such as the Stiefel, symplectic Stiefel, and Grassmann manifolds. By exploiting a global tangent space representation (the Lie subspace), the authors eliminate the need for projection steps and enable all Adam operations to be performed directly on these manifolds. The new optimizer is applied to train transformers and a symplectic autoencoder, achieving orthogonality constraints to machine precision and outperforming existing methods.
By Benedikt Brantner
arXiv:2510. 21033v3 Announce Type: replace-cross Abstract: We develop a theory of iso-Riemannian optimization for problems constrained to learned data manifolds, a setting in which classical Riemannian optimization - and Riemannian gradient descent in particular - can be poorly suited.
By Willem Diepeveen, Melanie Weber
The paper extends the analysis of Joint-Embedding Predictive Architectures (JEPAs) beyond Euclidean latent spaces to Riemannian manifolds. It shows that when latent variables lie on a sphere and the target distribution matches this spherical geometry, every optimal representation recovers the latent state up to an orthogonal transformation, demonstrating that Gaussian uniqueness is not universal. Experiments confirm that geometrically compatible targets improve linear recovery, especially in high-dimensional toroidal settings.
By L\'eo Nicollier (CB, ATT), Enric Meinhardt-Llopis (CB), Marc Pic (ATT), Pablo Mus\'e (CB, IFUMI), Gabriele Facciolo (CB)