arXiv Machine Learning

GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training

The paper investigates how post‑training modifies the weights of Large Language Models relative to their pretrained state. By expressing weight updates in the pretrained matrix’s singular value decomposition, the authors separate changes into three geometric components: diagonal (reshaping singular values), off‑diagonal (rotating input‑output coupling), and null‑space (routing outside the original SVD core). Experiments on a math evaluation suite show that removing the diagonal component largely preserves post‑training gains, indicating that improvements stem mainly from reconfiguring and extending pretrained pathways rather than altering singular values.

arXiv Machine Learning
Aug 12

Diffract: Spectral View of LLM Domain Adaptation

arXiv:2608. 10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text.

By Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko
arXiv Machine Learning
Jul 2

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

arXiv:2607. 01232v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers.

By Zijian Zhang, Rizhen Hu, Athanasios Glentis, Dawei Li, Chung-Yiu Yau, Hongzhou Lin, Mingyi Hong
arXiv AI
Jul 14

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

arXiv:2607. 11505v1 Announce Type: cross Abstract: Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment.

By Daocheng Fu, Rong Wu, Yu Yang, Xuemeng Yang, Jianbiao Mei, Licheng Wen, Pinlong Cai, Yong Liu, Botian Shi, Yu Qiao