arXiv Machine Learning

Nonlinearity-Aware LoRA: Structured Gate Adaptation under Low-Rank Constraints

arXiv:2606. 31717v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) is commonly viewed as an update-space approximation to full fine-tuning, yet this view is incomplete for self-gated Transformer feed-forward networks.

arXiv Machine Learning
Aug 31

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

The paper introduces Mixture of Activations (MoA), a token‑adaptive feedforward network design that mixes multiple activation functions using lightweight gates while sharing linear projections. It also presents learnable activations (LA) as an input‑independent variant. The authors theoretically prove that MoA strictly surpasses both fixed‑activation FFNs and LA in expressive power, and empirically demonstrate that MoA achieves lower loss and better scaling on dense and MoE language models from 0.12 B to 2 B parameters with minimal overhead.

By Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
arXiv AI
Sep 3

TaRA: Training-Aware Low-Rank Adaptation Initialization

TaRA: Training-Aware Low-Rank Adaptation Initialization proposes a new way to initialize LoRA by aligning the gradients of low‑rank factors with those of the full‑rank weight matrix. This approach directly incorporates training dynamics, improving gradient fidelity at the start of fine‑tuning while adding negligible computational cost. Experiments on a variety of challenging fine‑tuning tasks show that TaRA consistently outperforms existing state‑of‑the‑art initialization methods, offering a simple, robust, and scalable solution for effective LoRA initialization.

By Taehyeon Kim, Eunhyeok Park
arXiv Computer Vision
Aug 31

Activation Boundary Matching: Task-Informed Initialization for Low-Rank Adaptation

The paper introduces Activation Boundary Matching for Low‑Rank Adaptation (ABM‑LoRA), a task‑informed initialization strategy that uses the signs of layer‑wise pre‑activations from a brief probe adapter as targets for a fresh adapter. By training with a margin‑based hinge objective on these activation boundaries, ABM‑LoRA captures useful adaptation directions that standard LoRA initializers miss, while requiring only a few forward passes. Experiments show that ABM‑LoRA outperforms or matches existing LoRA, SVD, and gradient‑based initializers across multiple models and benchmarks, including T5‑base/GLUE, ConvNeXt‑T, Swin‑T, Qwen2.5‑1.5B, and LLaMA2‑7B.

By Dongha Lee, Jinhee Park, Minjun Kim, Junseok Kwon
arXiv Machine Learning
Jul 14

AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating

arXiv:2607. 10593v1 Announce Type: new Abstract: Normalization is a critical component for stabilizing Transformer training, yet the choice between static strategies such as Layer Normalization (LN) and adaptive alternatives remains largely task-dependent.

By Piyush Kaushik Bhattacharyya, Divyanshu Rai, Swastik Singh, Kumar Aakash, Ayush Ranjan, Krutika Verma