The paper introduces Hyper^2, a dual‑space consistency framework that applies hyperbolic geometry consistently to both the loss and the encoder in point‑cloud completion tasks. By reusing the same arcosh(1+αd²) function as a positional bias in refinement attention and as the Chamfer loss, Hyper^2 achieves significant Chamfer error reductions—up to 22.9% on ShapeNet‑55 and 37.5% on unseen ShapeNet‑34—while adding only ~1.6% FLOPs. The authors demonstrate that geometric consistency across encoder and loss, rather than either component alone, is key to effective hyperbolic supervision, supported by two model‑agnostic indicators that peak only when both are hyperbolic.
By Guantian Zheng, Haiyang Xu, Tianyu Gao
hyperbolix is an open‑source library for hyperbolic deep learning in JAX, built on Flax NNX. It provides six manifolds—including Euclidean, Poincaré ball, hyperboloid, κ‑stereographic, mixed‑curvature product, and proper velocity space—through a common interface, and implements a wide range of layer families (linear, convolution, attention, normalization, positional encoding, regression, vector quantization). The library also supplies Riemannian optimizers, wrapped distributions, dimensionality‑reduction techniques, and precision‑tested operations that replace numerically unstable formulas on the hyperboloid, ensuring accurate float32 computations at large distances.
By Timo Klein, Thomas Lang, Yllka Velaj, Sebastian Tschiatschek
arXiv:2608. 14803v1 Announce Type: new Abstract: A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift along the zero-loss manifold toward lower norm.
By Suvinava Basak
arXiv:2607. 05268v1 Announce Type: cross Abstract: Whether a hyperbolic representation model uses its geometry cannot be read off its curvature parameter: what matters is the dimensionless operating point $\sqrt{c}\rho$ and whether the radial and cone machinery is active there.
By Jaeyoung Kim, Eunseok Kim, Dongsuk Jang
arXiv:2608. 19584v1 Announce Type: new Abstract: We study landscapes for complex-parameterized networks.
By Andrew Gracyk
arXiv:2608. 02668v1 Announce Type: cross Abstract: Residual connections are the de facto mechanism for training deep neural networks stably.
By Jie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
arXiv:2606. 02596v1 Announce Type: new Abstract: The curvature exponent $\alpha$ in $h_k \propto \sigma_k^\alpha$ -- governing how Hessian eigenvalues scale with gradient singular values -- varies systematically across layer types ($\alpha \approx 2$ for convolutions, $\approx 1$ for transformer attention, $< 1$ for MLP up-projections).
By Anherutowa Calvo
arXiv:2609. 02672v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix.
By Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng, Ximen, Wenhan Luo
arXiv:2608. 05136v1 Announce Type: new Abstract: Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not.
By Devender Singh
arXiv:2607. 03998v1 Announce Type: new Abstract: The local sharpness of the loss, the top Hessian eigenvalue $\lambda_1$, determines the largest stable gradient step, but measuring it normally requires Lanczos or Hessian-vector iterations.
By Ashmitha R, J\"org Frochte
The paper introduces Orthogonal Hyper-Connections (oHC), a new approach that replaces the single residual stream of a Transformer with multiple parallel streams mixed by a rotation matrix from the group SO(n). By constraining the mixing matrix to SO(n) and parameterizing it with unit quaternions for four streams, oHC prevents both amplification and attenuation of residuals, maintaining training stability and preserving stream diversity. Experiments show that oHC outperforms the single-stream baseline, manifold-constrained Hyper-Connections, and identity-fixed Hyper-Connections across a wide range of downstream tasks.
The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.
By Rub\'en Dar\'io Guerrero