arXiv Machine Learning By Ravi Satya Durga Prasad Yenugula

Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't

Read the original on arXiv Machine Learning →

arXiv:2608. 02829v1 Announce Type: new Abstract: Model families train every size from scratch.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.

By Ravi Satya Durga Prasad Yenugula
arXiv Machine Learning
Sep 3

Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

The paper introduces a new approach to low‑rank clone distillation that ensures the student model’s training targets exactly the weights it will use at inference. By redefining the training objective to cover the full deployed matrix—without changing the model’s shape, parameter count, or FLOPs—the authors recover previously unreachable linear degrees of freedom. This results in significant performance gains across multiple teacher models, achieving comparable or superior accuracy with fewer tokens and parameters.

By Wenhui Chen, Zhifeng Li, Jie Zhou, Navan Preet Singh, Madalina Ciobanu, Chenghua Wang, Qingqing Mao, Ritankar Das
arXiv AI
Sep 11

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan