arXiv Machine Learning By Sen Song

Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design

Read the original on arXiv Machine Learning →

arXiv:2608. 01283v1 Announce Type: new Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices

The paper studies matrices built from block‑diagonal factors interleaved with fixed permutations, a structured family useful in deep learning for balancing expressivity and efficiency. By applying Riemannian geometry, the authors determine when this class forms a smooth manifold and develop Riemannian tools for the orthogonal two‑factor case. They propose efficient algorithms that use automatic differentiation, allow parameter sharing, and avoid dense matrix construction, testing them on matrix approximation and fine‑tuning large language models, while also exploring properties of factorizations with more factors.

By Ali Aliev, Maxim Rakhuba