arXiv:2607. 22757v1 Announce Type: cross Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective.
By T. Shaska
arXiv:2608. 28150v1 Announce Type: new Abstract: Which geometry controls the rank complexity of normalized softmax attention?
By Yuhe Sui, Jianing Zhang
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.
By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li
The paper introduces a computational tropical geometry framework for symbolically analyzing neural networks with tropical activations. It presents an algorithm that computes the network’s linear regions as explicit unions of polyhedra, proves its correctness, and connects the number of linear regions to the monomials in the tropical expression. The authors also define the Hoffman constant to bound distances to the farthest linear region and release the open‑source Julia library TropicalNN.jl to implement these tools, demonstrating their use on proof‑of‑concept examples.
By Paul Lezeau, Thomas Walker, Yueqi Cao, Shiv Bhatia, Anthea Monod
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie
arXiv:2602. 18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation.
By Seyed Morteza Emadi