The paper establishes rigorous trainability results for multi-headed attention layers and Low Rank Adaptation (LoRA) models under stochastic training methods. By proving that the empirical regression loss induces a Poincaré inequality with constants independent of data dimension for LoRA and independent of head dimensions for multi-head attention, the authors show that a stochastic differential equation mimicking SGD converges to the loss minima. These results hold without assumptions on data or model size, providing the first theoretical guarantees for training such architectures.
By Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun
arXiv:2607. 13425v1 Announce Type: cross Abstract: Learning effectively from limited data is critical in domains like security where labeled examples are scarce.
By Tuomas Oikarinen, Zixiao Chen, Charlotte Siska, Tsui-Wei Weng, Chandan Singh, Jianfeng Gao
The paper introduces Higher-Order Modular Attention (HOMA), a new attention mechanism that combines standard pairwise self‑attention with an explicit triadic attention pathway. HOMA uses overlapping blocks, local windows, and a low‑rank projection to make triadic interactions tractable. Experiments on controlled PARITY and MATCH3 tasks, as well as TAPE benchmarks, show that HOMA matches or outperforms matched pairwise and purely triadic baselines, especially when dependencies extend beyond triadic order, and it often converges faster and uses parameters more efficiently.
By Shirin Amiraslani, Xin Gao
arXiv:2508. 17821v3 Announce Type: replace-cross Abstract: This paper investigates the limitations of the normalization in attention mechanisms.
By Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State
arXiv:2606. 18587v1 Announce Type: cross Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens.
By Zhiyuan Wang, Xuan Luo, Sirui Zeng, Xifeng Yan
arXiv:2609.05885v1 Announce Type: new
Abstract: Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR...
By Huiyi Wang, Daijiao Liu, Lina Yao, Dong Gong