The paper establishes rigorous trainability results for multi-headed attention layers and Low Rank Adaptation (LoRA) models under stochastic training methods. By proving that the empirical regression loss induces a Poincaré inequality with constants independent of data dimension for LoRA and independent of head dimensions for multi-head attention, the authors show that a stochastic differential equation mimicking SGD converges to the loss minima. These results hold without assumptions on data or model size, providing the first theoretical guarantees for training such architectures.
By Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun
arXiv:2607. 13425v1 Announce Type: cross Abstract: Learning effectively from limited data is critical in domains like security where labeled examples are scarce.
By Tuomas Oikarinen, Zixiao Chen, Charlotte Siska, Tsui-Wei Weng, Chandan Singh, Jianfeng Gao
The paper introduces Higher-Order Modular Attention (HOMA), a new attention mechanism that combines standard pairwise self‑attention with an explicit triadic attention pathway. HOMA uses overlapping blocks, local windows, and a low‑rank projection to make triadic interactions tractable. Experiments on controlled PARITY and MATCH3 tasks, as well as TAPE benchmarks, show that HOMA matches or outperforms matched pairwise and purely triadic baselines, especially when dependencies extend beyond triadic order, and it often converges faster and uses parameters more efficiently.
By Shirin Amiraslani, Xin Gao
arXiv:2508. 17821v3 Announce Type: replace-cross Abstract: This paper investigates the limitations of the normalization in attention mechanisms.
By Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State
arXiv:2606. 18587v1 Announce Type: cross Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens.
By Zhiyuan Wang, Xuan Luo, Sirui Zeng, Xifeng Yan
arXiv:2609.05885v1 Announce Type: new
Abstract: Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR...
By Huiyi Wang, Daijiao Liu, Lina Yao, Dong Gong
arXiv:2505. 15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic.
By Suvadeep Hajra
arXiv:2509. 07963v2 Announce Type: replace Abstract: The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product of each pair.
By Yilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu, Andres Potapczynski, Andrew Gordon Wilson
arXiv:2605.07111v3 Announce Type: replace-cross
Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater represe...
By Haozhan Tang, Xiuqi Zhu, Xinyin Zhang, Boxun Li, Virginia Smith, Kevin Kuo
The paper introduces Activation Boundary Matching for Low‑Rank Adaptation (ABM‑LoRA), a task‑informed initialization strategy that uses the signs of layer‑wise pre‑activations from a brief probe adapter as targets for a fresh adapter. By training with a margin‑based hinge objective on these activation boundaries, ABM‑LoRA captures useful adaptation directions that standard LoRA initializers miss, while requiring only a few forward passes. Experiments show that ABM‑LoRA outperforms or matches existing LoRA, SVD, and gradient‑based initializers across multiple models and benchmarks, including T5‑base/GLUE, ConvNeXt‑T, Swin‑T, Qwen2.5‑1.5B, and LLaMA2‑7B.
By Dongha Lee, Jinhee Park, Minjun Kim, Junseok Kwon
arXiv:2603. 03993v2 Announce Type: replace Abstract: Multi-head attention enables transformer models to represent multiple attention patterns simultaneously.
By M. Sagitova, O. Duranthon, L. Zdeborov\'a
arXiv:2609.08615v1 Announce Type: new
Abstract: Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single sca...
By Lenard Dome