arXiv:2608. 02668v1 Announce Type: cross Abstract: Residual connections are the de facto mechanism for training deep neural networks stably.
By Jie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
arXiv:2608. 10416v1 Announce Type: cross Abstract: We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realization (Riemann GeoResolver).
By Liangchen Ge
arXiv:2510. 21033v3 Announce Type: replace-cross Abstract: We develop a theory of iso-Riemannian optimization for problems constrained to learned data manifolds, a setting in which classical Riemannian optimization - and Riemannian gradient descent in particular - can be poorly suited.
By Willem Diepeveen, Melanie Weber
arXiv:2608. 01283v1 Announce Type: new Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al.
By Sen Song
Weight-space geometry plays a central role in neural network optimization, yet manifold constraints are often applied uniformly across all weight matrices. In this work, we ask whether different transformer modules prefer different manifold geometries.
arXiv:2506. 21278v3 Announce Type: replace-cross Abstract: We propose spherical Cauchy (spCauchy) latent variables for variational autoencoders on hyperspherical latent spaces.
By Lukas Sablica, Kurt Hornik