arXiv Machine Learning

Low-Rank Attention Residuals

arXiv:2607. 09694v1 Announce Type: new Abstract: Attention Residuals replace the fixed residual sum with depthwise attention over previous sub-layer outputs in large language models (LLMs), but use each output as both a full-dimensional key and value.

arXiv AI
6d ago

Attention Sinks and Outliers in Attention Residuals

The paper introduces OASIS, a method designed to stabilize dual‑normalized attention‑residual architectures by employing explicit null routing and token‑to‑depth null coupling. OASIS mitigates attention sinks and activation outliers, improving low‑bit quantization performance across several language‑model backbones. Empirical results show significant reductions in attention norms and perplexity, with notable gains on long‑context benchmarks.

By Haozheng Luo, Haoran Dai, Jingyuan Huang, Shaoyang Zhang, Xi Chen, Eric Hanchen Jiang, Yijiang Li, Chenghao Qiu, Chenwei Xu, Zhenyu Pan, Haotian Zhang, Binghui Wang, Yan Chen
arXiv AI
Sep 21

Attention-Aware Routing: Coupling Routing and Attention in MoEs

Attention-Aware Routing (AAR) augments the router in Mixture-of-Experts language models with temporal and spectral features derived from a sliding window of attention weights, thereby separating contextual information from the token’s hidden state. By keeping the base transformer frozen and training only routing parameters, AAR achieves a +3.37‑point improvement on GSM8K over a routing‑only baseline and demonstrates that routing changes propagate through the residual stream to reshape attention without directly updating the attention mechanism. The method also reduces long diverging generations, shows depth‑sensitivity affecting retrieval versus reasoning, and offers a controlled probe of routing‑relevant information across layers.

By Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou, Giannis Karamanolakis, Swastik Roy, Alexandros Potamianos
arXiv AI
Jul 8

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

arXiv:2607. 06523v1 Announce Type: new Abstract: Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply uniform budgets across layers or tokens and degrade retrieval when lexical cues and semantic states require different preservation.

By Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jesus Olivera
arXiv AI
Sep 11

Forward-Free LLM Depth Pruning via Weight Redundancy

The paper introduces Weight-Redundancy Pruning (WRP), a forward‑free depth‑pruning technique for large language models that estimates inter‑layer redundancy using only checkpoint weights. WRP compares attention outputs and MLP down‑projection weights across layers, combining pairwise similarities with relative projection‑scale information to guide layer grouping and block selection. Experiments show that WRP consistently outperforms existing forward‑free magnitude pruning methods and approaches the performance of activation‑based pruning across various pruning settings, model families, and downstream tasks.

By Vincent-Daniel Yun, Woosang Lim
arXiv Machine Learning
4d ago

On State Reduction in Linear Attention

arXiv:2602.04852v3 Announce Type: replace Abstract: Linear attention offers a computationally efficient yet expressive alternative to softmax attention. However, recent empirical results indicate tha...

By Philipp Nazari, T. Konstantin Rusch
arXiv Computer Vision
Aug 27

SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs

SHIFT-LLM is a training‑free post‑pruning correction framework that inserts a Linear Residual Adapter (LRA) at each depth‑pruned site in large language models. Each LRA preserves the original residual identity while adding a lightweight affine correction calibrated via closed‑form least‑squares regression on a small held‑out set, thereby approximating the hidden state that would have been produced by the removed block. Experiments across multiple model families and benchmarks show that SHIFT‑LLM consistently recovers accuracy lost to depth pruning, achieving gains up to +15.7 points on Llama‑3.1‑8B‑Instruct with only a few hundred calibration samples and no gradient computation.

By Ali Bahri, Hang Li, Hongliang Li, Zhitang Chen