arXiv AI

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.

arXiv Machine Learning
Sep 3

Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

The paper introduces a new approach to low‑rank clone distillation that ensures the student model’s training targets exactly the weights it will use at inference. By redefining the training objective to cover the full deployed matrix—without changing the model’s shape, parameter count, or FLOPs—the authors recover previously unreachable linear degrees of freedom. This results in significant performance gains across multiple teacher models, achieving comparable or superior accuracy with fewer tokens and parameters.

By Wenhui Chen, Zhifeng Li, Jie Zhou, Navan Preet Singh, Madalina Ciobanu, Chenghua Wang, Qingqing Mao, Ritankar Das
arXiv AI
Sep 11

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv AI
Sep 10

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.

By Yunpeng Xu, Kun Zheng
arXiv Machine Learning
Sep 17

Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability

The study investigates which components of a neural network contribute to rapid generalization (grokking) and how stable that improvement remains during further training. By transferring internal attention and MLP weights along with token embeddings and readout, the authors achieve a 5.46‑percentage‑point boost in early accuracy and a 558‑step reduction in confirmation latency, while also demonstrating that freezing transferred representations largely prevents post‑grokking relapse. The work delineates clear component‑level differences between acceleration and stability, and identifies architectural limits where omitting donor embeddings leads to significant performance loss.

By Zeyu Jia
Hugging Face Trending Papers
Sep 8

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The paper investigates how per-domain data composition during the mid‑training phase (between pre‑training and alignment) affects model performance. Experiments with Qwen3‑8B‑Base across five KOR‑Bench domains show that a moderate coverage band (10%‑40%) yields the best performance for each domain, and that alignment passes cannot fully close the gaps created by suboptimal mid‑training allocations. Additionally, zero coverage during mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.