arXiv AI

Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression

The paper introduces a scalable Kronecker-based approximation that captures cross-layer interactions without storing the full Fisher matrix, making Hessian analysis feasible for billion-parameter language models. It identifies consistent vulnerability patterns, notably that value projection layers are the most sensitive and exhibit strong cross-layer correlations across various model families. Experiments on quantization, sparsification, inter-layer corruption, and fine-tuning show that the approximation correlates strongly with performance degradation and recovery, providing a practical tool for identifying fragile components and guiding compression and optimization strategies.

arXiv Machine Learning
Sep 4

MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

The paper introduces MSign, an optimizer designed to prevent training instability in large language models by restoring the stable rank of weight matrices. It identifies two precursors to gradient explosions—rapid stable rank decline and increased Jacobian alignment—and proves that these jointly cause exponential gradient growth. Experiments on models ranging from 5 M to 3 B parameters show that MSign stops training failures while adding less than 7.0% computational overhead.

By Lianhai Ren, Yucheng Ding, Xiao Liu, Peng Cheng, Yeyun Gong
arXiv Machine Learning
Sep 10

Dense Structural Compression of Transformers via Gauge-Correct Channel Removal

The paper introduces GaugeLasso, a method that applies symmetric group‑lasso penalties to transformer channels during training, enabling entire tensor slices to be zeroed out while maintaining dense tensors for GPU efficiency. By calibrating channel penalties based on inference utility per compute, the network self‑organizes into depth‑dependent structural profiles that can be dramatically smaller than the original architecture, achieving up to 255‑fold compression on a polynomial division task and outperforming hand‑designed baselines on language modeling and autoencoding benchmarks. The approach also accelerates training and reveals over‑provisioned axes that guide subsequent design iterations.

By Jed A. Duersch, Na\"im Es-Sebbani, Nathana\"el Haas, Zied Bouraoui
arXiv AI
Sep 1

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.

By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki