arXiv:2607. 22757v1 Announce Type: cross Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective.
By T. Shaska
arXiv:2608. 28150v1 Announce Type: new Abstract: Which geometry controls the rank complexity of normalized softmax attention?
By Yuhe Sui, Jianing Zhang
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.
By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li
The paper introduces a computational tropical geometry framework for symbolically analyzing neural networks with tropical activations. It presents an algorithm that computes the network’s linear regions as explicit unions of polyhedra, proves its correctness, and connects the number of linear regions to the monomials in the tropical expression. The authors also define the Hoffman constant to bound distances to the farthest linear region and release the open‑source Julia library TropicalNN.jl to implement these tools, demonstrating their use on proof‑of‑concept examples.
By Paul Lezeau, Thomas Walker, Yueqi Cao, Shiv Bhatia, Anthea Monod
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie
arXiv:2602. 18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation.
By Seyed Morteza Emadi
arXiv:2609. 04046v1 Announce Type: cross Abstract: What can a single layer of self-attention compute?
By Rajmohan Rajaraman, Ravi Sundaram, Amanuel Tesfaye
arXiv:2609.21523v1 Announce Type: new
Abstract: A system may be compressed before its downstream task is fully known. We ask how much retained state is then necessary and how much can be saved by lim...
By Ronald Katende
arXiv:2607. 10677v1 Announce Type: new Abstract: Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood.
By Binbin Lin, Wei Chen, Yalun Li, Wenxiao Wang, Jieping Ye, Xiaofei He
arXiv:2609. 18145v1 Announce Type: new Abstract: Attention pays, at every layer and for every input, the cost of searching for whom to connect.
By Yoshiaki Takashita
arXiv:2605. 28983v2 Announce Type: replace-cross Abstract: In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights.
By Jose Marie Antonio Mi\~noza, Erika Fille T. Legara, Christopher P. Monterola
arXiv:2609.01129v1 Announce Type: new
Abstract: We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under compos...
By Jiming Feng, Junliang Li