arXiv AI

After the Euclidean Highway: Hyperbolic Expert AI as the Next Innovation

arXiv:2607. 17513v1 Announce Type: cross Abstract: Expert domains are trees; the Euclidean transformer is not, diluting parent-child structure exponentially at depth.

arXiv Computer Vision
Aug 25

Hyper^2: Unleashing Hyperbolic Geometry's Full Potential via Dual-Space Consistency

The paper introduces Hyper^2, a dual‑space consistency framework that applies hyperbolic geometry consistently to both the loss and the encoder in point‑cloud completion tasks. By reusing the same arcosh(1+αd²) function as a positional bias in refinement attention and as the Chamfer loss, Hyper^2 achieves significant Chamfer error reductions—up to 22.9% on ShapeNet‑55 and 37.5% on unseen ShapeNet‑34—while adding only ~1.6% FLOPs. The authors demonstrate that geometric consistency across encoder and loss, rather than either component alone, is key to effective hyperbolic supervision, supported by two model‑agnostic indicators that peak only when both are hyperbolic.

By Guantian Zheng, Haiyang Xu, Tianyu Gao
arXiv Machine Learning
Sep 24

hyperbolix: Hyperbolic Deep Learning in JAX

hyperbolix is an open‑source library for hyperbolic deep learning in JAX, built on Flax NNX. It provides six manifolds—including Euclidean, Poincaré ball, hyperboloid, κ‑stereographic, mixed‑curvature product, and proper velocity space—through a common interface, and implements a wide range of layer families (linear, convolution, attention, normalization, positional encoding, regression, vector quantization). The library also supplies Riemannian optimizers, wrapped distributions, dimensionality‑reduction techniques, and precision‑tested operations that replace numerically unstable formulas on the hyperboloid, ensuring accurate float32 computations at large distances.

By Timo Klein, Thomas Lang, Yllka Velaj, Sebastian Tschiatschek
arXiv AI
Aug 5

Sphere Retraction Normalizations

arXiv:2608. 02668v1 Announce Type: cross Abstract: Residual connections are the de facto mechanism for training deep neural networks stably.

By Jie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
arXiv Machine Learning
Jun 3

Spectral Asymptotics of Neural Network Loss Landscapes: An Exact Decomposition of the Curvature Exponent

arXiv:2606. 02596v1 Announce Type: new Abstract: The curvature exponent $\alpha$ in $h_k \propto \sigma_k^\alpha$ -- governing how Hessian eigenvalues scale with gradient singular values -- varies systematically across layer types ($\alpha \approx 2$ for convolutions, $\approx 1$ for transformer attention, $< 1$ for MLP up-projections).

By Anherutowa Calvo
arXiv Machine Learning
Sep 3

oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions

arXiv:2609. 02672v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix.

By Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng, Ximen, Wenhan Luo
Hugging Face Trending Papers
Sep 2

oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions

The paper introduces Orthogonal Hyper-Connections (oHC), a new approach that replaces the single residual stream of a Transformer with multiple parallel streams mixed by a rotation matrix from the group SO(n). By constraining the mixing matrix to SO(n) and parameterizing it with unit quaternions for four streams, oHC prevents both amplification and attenuation of residuals, maintaining training stability and preserving stream diversity. Experiments show that oHC outperforms the single-stream baseline, manifold-constrained Hyper-Connections, and identity-fixed Hyper-Connections across a wide range of downstream tasks.

arXiv Machine Learning
Sep 18

Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not

The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.

By Rub\'en Dar\'io Guerrero