arXiv AI By Kaivan Kamali, Kajetan Schweighofer, Hormoz Shahrzad, Olivier Francon, Babak Hodjat, Risto Miikkulainen

Efficient Pre-Training of LLMs through Truncated SVD Representations

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Sep 22

GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training

The paper investigates how post‑training modifies the weights of Large Language Models relative to their pretrained state. By expressing weight updates in the pretrained matrix’s singular value decomposition, the authors separate changes into three geometric components: diagonal (reshaping singular values), off‑diagonal (rotating input‑output coupling), and null‑space (routing outside the original SVD core). Experiments on a math evaluation suite show that removing the diagonal component largely preserves post‑training gains, indicating that improvements stem mainly from reconfiguring and extending pretrained pathways rather than altering singular values.

By Jianing Qi, Hao Tang, Zhigang Zhu
arXiv Machine Learning
Jul 17

Stabilizing Native Low-Rank LLM Pretraining

arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.

By Paul Janson, Edouard Oyallon, Eugene Belilovsky