hyperbolix is an open‑source library for hyperbolic deep learning in JAX, built on Flax NNX. It provides six manifolds—including Euclidean, Poincaré ball, hyperboloid, κ‑stereographic, mixed‑curvature product, and proper velocity space—through a common interface, and implements a wide range of layer families (linear, convolution, attention, normalization, positional encoding, regression, vector quantization). The library also supplies Riemannian optimizers, wrapped distributions, dimensionality‑reduction techniques, and precision‑tested operations that replace numerically unstable formulas on the hyperboloid, ensuring accurate float32 computations at large distances.
By Timo Klein, Thomas Lang, Yllka Velaj, Sebastian Tschiatschek
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...
arXiv:2608. 10416v1 Announce Type: cross Abstract: We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realization (Riemann GeoResolver).
By Liangchen Ge
The paper introduces Hyper^2, a dual‑space consistency framework that applies hyperbolic geometry consistently to both the loss and the encoder in point‑cloud completion tasks. By reusing the same arcosh(1+αd²) function as a positional bias in refinement attention and as the Chamfer loss, Hyper^2 achieves significant Chamfer error reductions—up to 22.9% on ShapeNet‑55 and 37.5% on unseen ShapeNet‑34—while adding only ~1.6% FLOPs. The authors demonstrate that geometric consistency across encoder and loss, rather than either component alone, is key to effective hyperbolic supervision, supported by two model‑agnostic indicators that peak only when both are hyperbolic.
By Guantian Zheng, Haiyang Xu, Tianyu Gao
arXiv:2609.27988v1 Announce Type: cross
Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...
By Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
The paper introduces the spherical Cauchy distribution as a new hyperspherical posterior for variational autoencoders, avoiding the complications of the von Mises–Fisher and Power Spherical alternatives. By using stereographic projection and a Möbius transformation, the authors obtain exact posterior samples and a closed‑form KL divergence that terminates in a finite polynomial for even dimensions and admits certified truncation for odd dimensions. Empirical results show that the spherical Cauchy yields faster inference and lower reconstruction loss on MNIST and improved negative log‑likelihood on smallNORB compared to existing methods.
By Lukas Sablica, Kurt Hornik
arXiv:2609.37817v1 Announce Type: new
Abstract: Geometric representation learning predominantly scaffolds representations onto flat Euclidean subspaces or compact product tori ($\mathbb{T}^K$). Howev...
By Zhongping Ji
The paper extends the analysis of Joint-Embedding Predictive Architectures (JEPAs) beyond Euclidean latent spaces to Riemannian manifolds. It shows that when latent variables lie on a sphere and the target distribution matches this spherical geometry, every optimal representation recovers the latent state up to an orthogonal transformation, demonstrating that Gaussian uniqueness is not universal. Experiments confirm that geometrically compatible targets improve linear recovery, especially in high-dimensional toroidal settings.
By L\'eo Nicollier (CB, ATT), Enric Meinhardt-Llopis (CB), Marc Pic (ATT), Pablo Mus\'e (CB, IFUMI), Gabriele Facciolo (CB)
A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation. Adam's per-coordinate preconditioner drifts along each symmetry orbit, which pulls the trajectory off the symmetry quotient where the optimization lives and blurs the singular-learning rate the quotient makes readable.
The paper introduces Orthogonal Hyper-Connections (oHC), a new approach that replaces the single residual stream of a Transformer with multiple parallel streams mixed by a rotation matrix from the group SO(n). By constraining the mixing matrix to SO(n) and parameterizing it with unit quaternions for four streams, oHC prevents both amplification and attenuation of residuals, maintaining training stability and preserving stream diversity. Experiments show that oHC outperforms the single-stream baseline, manifold-constrained Hyper-Connections, and identity-fixed Hyper-Connections across a wide range of downstream tasks.
arXiv:2606. 29176v1 Announce Type: new Abstract: A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation.
By Tejas Pradeep Shirodkar
arXiv:2608. 14803v1 Announce Type: new Abstract: A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift along the zero-loss manifold toward lower norm.
By Suvinava Basak