arXiv Machine Learning

Complexity-Guided Component-wise Initialization for Language Model Pretraining

arXiv:2607. 09204v1 Announce Type: cross Abstract: Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization.

arXiv AI
Jul 7

Spectral Signatures of Large Language Models

arXiv:2607. 03377v1 Announce Type: cross Abstract: The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management and quantification at scale, such as model lineage tracing, licensing, and evaluation.

By Zhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu, Hengrui Luo, Pu Ren, Yaoqing Yang
arXiv Machine Learning
Jul 17

Stabilizing Native Low-Rank LLM Pretraining

arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.

By Paul Janson, Edouard Oyallon, Eugene Belilovsky
arXiv Machine Learning
Aug 12

Diffract: Spectral View of LLM Domain Adaptation

arXiv:2608. 10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text.

By Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko
arXiv Computation and Language
Sep 4

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.

By Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang
arXiv Machine Learning
Aug 5

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

arXiv:2608. 03494v1 Announce Type: cross Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency.

By Raviraj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan, Niranjan Wartikar
Hugging Face Trending Papers
Sep 3

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.

arXiv AI
Jul 28

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.

By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv Machine Learning
Sep 4

MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

The paper introduces MSign, an optimizer designed to prevent training instability in large language models by restoring the stable rank of weight matrices. It identifies two precursors to gradient explosions—rapid stable rank decline and increased Jacobian alignment—and proves that these jointly cause exponential gradient growth. Experiments on models ranging from 5 M to 3 B parameters show that MSign stops training failures while adding less than 7.0% computational overhead.

By Lianhai Ren, Yucheng Ding, Xiao Liu, Peng Cheng, Yeyun Gong
arXiv Machine Learning
1d ago

Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining

The study investigates how layer‑wise intervention responses in language models change over the course of pretraining, using single‑block identity bypass across multiple checkpoints and model‑domain combinations. It finds that while depth ordering of responses persists, their magnitudes shift, with nearby checkpoints showing stronger rank correspondence than distant ones and large changes occurring at positions that recur across samples and transfer across evaluation domains. Controlled experiments reveal that these longitudinal changes cannot be explained by a single downstream sensitivity and depend on perturbation strength and direction, indicating that layer sensitivity is structured but dynamic.

By Shengye Tao, Yinzhu Cheng, Haihua Xie