Efficient Pre-Training of LLMs through Truncated SVD Representations
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 07098v1 Announce Type: cross Abstract: We present SigmaScale, a method for learning auxiliary scaling matrices $S$ to aid truncated Singular Value Decomposition (SVD) based Large Language Model (LLM) compression.
arXiv:2602.02848v2 Announce Type: replace Abstract: Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD...
The paper investigates how post‑training modifies the weights of Large Language Models relative to their pretrained state. By expressing weight updates in the pretrained matrix’s singular value decomposition, the authors separate changes into three geometric components: diagonal (reshaping singular values), off‑diagonal (rotating input‑output coupling), and null‑space (routing outside the original SVD core). Experiments on a math evaluation suite show that removing the diagonal component largely preserves post‑training gains, indicating that improvements stem mainly from reconfiguring and extending pretrained pathways rather than altering singular values.
arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.
arXiv:2606. 06494v1 Announce Type: new Abstract: Parameter-efficient finetuning methods based on spectral decomposition have enabled progress in Continual Learning.
arXiv:2606. 19993v1 Announce Type: new Abstract: We present Activation- and Influence-Aware Ranks (AIR), an SVD-based LLM compression framework that guides each weight matrix's low-rank approximation with a backward-signal influence metric.