arXiv Machine Learning By Zeyu Jia

Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability

Read the original on arXiv Machine Learning →

The study investigates which components of a neural network contribute to rapid generalization (grokking) and how stable that improvement remains during further training. By transferring internal attention and MLP weights along with token embeddings and readout, the authors achieve a 5.46‑percentage‑point boost in early accuracy and a 558‑step reduction in confirmation latency, while also demonstrating that freezing transferred representations largely prevents post‑grokking relapse. The work delineates clear component‑level differences between acceleration and stability, and identifies architectural limits where omitting donor embeddings leads to significant performance loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.

By Ravi Satya Durga Prasad Yenugula