Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
arXiv:2608. 02829v1 Announce Type: new Abstract: Model families train every size from scratch.
The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.
arXiv:2608. 02829v1 Announce Type: new Abstract: Model families train every size from scratch.
The paper introduces a new approach to low‑rank clone distillation that ensures the student model’s training targets exactly the weights it will use at inference. By redefining the training objective to cover the full deployed matrix—without changing the model’s shape, parameter count, or FLOPs—the authors recover previously unreachable linear degrees of freedom. This results in significant performance gains across multiple teacher models, achieving comparable or superior accuracy with fewer tokens and parameters.
The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.
The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.
arXiv:2607. 15525v1 Announce Type: cross Abstract: Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks.
arXiv:2609.14715v1 Announce Type: new Abstract: We scale our conventional sub-150M pretraining recipe from 53.5M to 109.7M parameters, holding the method fixed (Qwen3-style decoder with grouped-query...
The study investigates which components of a neural network contribute to rapid generalization (grokking) and how stable that improvement remains during further training. By transferring internal attention and MLP weights along with token embeddings and readout, the authors achieve a 5.46‑percentage‑point boost in early accuracy and a 558‑step reduction in confirmation latency, while also demonstrating that freezing transferred representations largely prevents post‑grokking relapse. The work delineates clear component‑level differences between acceleration and stability, and identifies architectural limits where omitting donor embeddings leads to significant performance loss.
arXiv:2605. 18838v3 Announce Type: replace-cross Abstract: Scaling laws predict loss from compute but not how capabilities interact.
arXiv:2606. 02608v1 Announce Type: new Abstract: We study a Marchenko--Pastur (MP) random-matrix approach to pruning deep neural networks with very small post-pruning fine-tuning budgets.
arXiv:2607. 14516v1 Announce Type: new Abstract: Interpreting optimizers as gradient-flow discretizations has motivated applying higher-order Runge-Kutta (RK) integrators to neural networks.
arXiv:2607. 13380v1 Announce Type: new Abstract: Predictive Coding (PC) offers a biologically motivated alternative to backpropagation via local weight updates, yet routing error between layers still relies on an autograd Jacobian-transpose ($J^\top$) product - the last non-local operation in PC.
The paper investigates how per-domain data composition during the mid‑training phase (between pre‑training and alignment) affects model performance. Experiments with Qwen3‑8B‑Base across five KOR‑Bench domains show that a moderate coverage band (10%‑40%) yields the best performance for each domain, and that alignment passes cannot fully close the gaps created by suboptimal mid‑training allocations. Additionally, zero coverage during mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.