arXiv Machine Learning

High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model

arXiv:2606. 05899v1 Announce Type: new Abstract: We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning.

arXiv Machine Learning
Sep 10

Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA

The paper establishes rigorous trainability results for multi-headed attention layers and Low Rank Adaptation (LoRA) models under stochastic training methods. By proving that the empirical regression loss induces a Poincaré inequality with constants independent of data dimension for LoRA and independent of head dimensions for multi-head attention, the authors show that a stochastic differential equation mimicking SGD converges to the loss minima. These results hold without assumptions on data or model size, providing the first theoretical guarantees for training such architectures.

By Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun
arXiv Machine Learning
Sep 25

Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning

The paper introduces Higher-Order Modular Attention (HOMA), a new attention mechanism that combines standard pairwise self‑attention with an explicit triadic attention pathway. HOMA uses overlapping blocks, local windows, and a low‑rank projection to make triadic interactions tractable. Experiments on controlled PARITY and MATCH3 tasks, as well as TAPE benchmarks, show that HOMA matches or outperforms matched pairwise and purely triadic baselines, especially when dependencies extend beyond triadic order, and it often converges faster and uses parameters more efficiently.

By Shirin Amiraslani, Xin Gao
arXiv Computer Vision
Aug 31

Activation Boundary Matching: Task-Informed Initialization for Low-Rank Adaptation

The paper introduces Activation Boundary Matching for Low‑Rank Adaptation (ABM‑LoRA), a task‑informed initialization strategy that uses the signs of layer‑wise pre‑activations from a brief probe adapter as targets for a fresh adapter. By training with a margin‑based hinge objective on these activation boundaries, ABM‑LoRA captures useful adaptation directions that standard LoRA initializers miss, while requiring only a few forward passes. Experiments show that ABM‑LoRA outperforms or matches existing LoRA, SVD, and gradient‑based initializers across multiple models and benchmarks, including T5‑base/GLUE, ConvNeXt‑T, Swin‑T, Qwen2.5‑1.5B, and LLaMA2‑7B.

By Dongha Lee, Jinhee Park, Minjun Kim, Junseok Kwon