The paper establishes rigorous trainability results for multi-headed attention layers and Low Rank Adaptation (LoRA) models under stochastic training methods. By proving that the empirical regression loss induces a Poincaré inequality with constants independent of data dimension for LoRA and independent of head dimensions for multi-head attention, the authors show that a stochastic differential equation mimicking SGD converges to the loss minima. These results hold without assumptions on data or model size, providing the first theoretical guarantees for training such architectures.
By Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun
arXiv:2604. 00757v2 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens.
By Dong-Jae Lee, Sunghyun Baek, Junmo Kim
arXiv:2601. 12145v3 Announce Type: replace Abstract: Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase.
By Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, Tong Zhao
arXiv:2509. 07963v2 Announce Type: replace Abstract: The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product of each pair.
By Yilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu, Andres Potapczynski, Andrew Gordon Wilson
arXiv:2607. 20214v1 Announce Type: cross Abstract: The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths.
By Mahdi Heidari, Mohammad Mahdi Rahimi, Jaekyun Moon
arXiv:2602.04852v3 Announce Type: replace
Abstract: Linear attention offers a computationally efficient yet expressive alternative to softmax attention. However, recent empirical results indicate tha...
By Philipp Nazari, T. Konstantin Rusch
arXiv:2608. 09417v2 Announce Type: replace Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes.
By Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang
arXiv:2610.08341v1 Announce Type: cross
Abstract: Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redu...
By Shuo Yang, Changbai Li, Linlin Yang, Huobin Tan, Rongyu Chen, Tongfei Chen, Tian Wang, Sheng Xu, Baochang Zhang
arXiv:2608. 09417v1 Announce Type: new Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes.
By Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang
arXiv:2604. 17324v2 Announce Type: replace-cross Abstract: Global self-attention drives modern graph transformers, yet the softmax at its core imposes a structural constraint rarely examined directly: every attention row is non-negative and sums to one, so each per-head output is a mass-conserving convex combination of value vectors.
By Yang Liu, Dongxin Guo, Tom Zheng, Siu Ming Yiu, Liam Ning, Jikun Wu
arXiv:2607. 23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it?
By Byeong Hoon Yoon
arXiv:2609.00667v1 Announce Type: cross
Abstract: Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning es...
By Siyi Liu, Hanjun Yang, Chenchen Zhang, Xiaorong Zhu, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang