arXiv Machine Learning

Diffract: Spectral View of LLM Domain Adaptation

arXiv:2608. 10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text.

arXiv Machine Learning
Jul 17

Stabilizing Native Low-Rank LLM Pretraining

arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.

By Paul Janson, Edouard Oyallon, Eugene Belilovsky
Hugging Face Trending Papers
Jun 3

STaR-Quant: State-Time Consistent Post-Training Quantization for Diffusion Large Language Models

Diffusion large language models (DLLMs) have recently emerged as a promising alternative to autoregressive LLMs by generating text through iterative masked denoising with bidirectional context. However, their large model sizes and iterative denoising process introduce substantial memory and computational overhead, motivating post-training quantization for efficient deployment.

arXiv Machine Learning
Sep 22

GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training

The paper investigates how post‑training modifies the weights of Large Language Models relative to their pretrained state. By expressing weight updates in the pretrained matrix’s singular value decomposition, the authors separate changes into three geometric components: diagonal (reshaping singular values), off‑diagonal (rotating input‑output coupling), and null‑space (routing outside the original SVD core). Experiments on a math evaluation suite show that removing the diagonal component largely preserves post‑training gains, indicating that improvements stem mainly from reconfiguring and extending pretrained pathways rather than altering singular values.

By Jianing Qi, Hao Tang, Zhigang Zhu
Hugging Face Trending Papers
Jun 9

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models

Diffusion large language models (dLLMs) re-encode the entire prefix at every denoising step, causing recomputation that scales quadratically with context length and becomes prohibitive for long-context scenarios. We propose Prefilling-dLLM, a training-free prefill-decode disaggregation framework for dLLMs that partitions the prefix into N chunks, caches their KV representations once, and selects the top-K most relevant chunks with intra-chunk token sparsity for decoding, showing that sparse prefilling can outperform dense attention while reducing per-step complexity from quadratic in the full sequence length to quadratic only in the decode length.