The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.
By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li
arXiv:2607. 20538v1 Announce Type: cross Abstract: Long-context Transformer inference increasingly relies on KV-cache compression or quantization.
By Yitao Jiang, Yaoqing Yang, Luyang Zhao, Muhao Chen, Devin Balkcom
arXiv:2607. 11883v1 Announce Type: new Abstract: Compression is fundamental to intelligence.
By Shikai Qiu, Marc Finzi, Yujia Zheng, Kun Zhang, Andrew Gordon Wilson
arXiv:2609.08851v1 Announce Type: new
Abstract: Recent advancements in transformer length generalization theory enable us to reliably predict when a transformer can learn to solve a task. In particul...
By Georg Zetzsche, Hongjian Jiang, Andy Yang, Pascal Bergstr\"a{\ss}er, Marco S\"alzer, David Chiang, Anthony W. Lin
arXiv:2505. 23869v4 Announce Type: replace-cross Abstract: A proposition that connects randomness and compression is put forward via Gibbs entropy over set of measurement vectors associated with a compression process.
By M. S\"uzen
arXiv:2411. 09816v5 Announce Type: replace Abstract: Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices.
By Cem \"Uy\"uk, Mike Lasby, Mohamed Yassin, Utku Evci, Yani Ioannou
Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization.
The paper introduces the concept of Information Capacity (IC) to quantify how much bandwidth savings a unit of decoder compute can achieve in generative video compression (GVC). By modeling reconstruction quality as a two‑factor power law in data rate and compute, the authors fit measured DISTS of two GVC decoders with high accuracy and define IC as the negative logarithmic slope along an iso‑quality contour. IC is dimensionless, enabling architecture‑agnostic comparisons and revealing that a 14B decoder trades compute for rate far more efficiently than a 1.3B decoder, with significant variation across datasets.
By Cheng Yuan, Jiawei Shao, Xuelong Li
arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.
By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv:2607. 20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams.
By Andrew Mack, Kraig Yuheng Tou, Mark Henry, Zhengxun Wu, Lauren Greenspan
arXiv:2609.15975v1 Announce Type: cross
Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
By Shwai He, Haichao Zhang, Shen Yan
The paper introduces GaugeLasso, a method that applies symmetric group‑lasso penalties to transformer channels during training, enabling entire tensor slices to be zeroed out while maintaining dense tensors for GPU efficiency. By calibrating channel penalties based on inference utility per compute, the network self‑organizes into depth‑dependent structural profiles that can be dramatically smaller than the original architecture, achieving up to 255‑fold compression on a polynomial division task and outperforming hand‑designed baselines on language modeling and autoencoding benchmarks. The approach also accelerates training and reveals over‑provisioned axes that guide subsequent design iterations.
By Jed A. Duersch, Na\"im Es-Sebbani, Nathana\"el Haas, Zied Bouraoui