Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability
Read the original on arXiv Machine Learning →The study investigates which components of a neural network contribute to rapid generalization (grokking) and how stable that improvement remains during further training. By transferring internal attention and MLP weights along with token embeddings and readout, the authors achieve a 5.46‑percentage‑point boost in early accuracy and a 558‑step reduction in confirmation latency, while also demonstrating that freezing transferred representations largely prevents post‑grokking relapse. The work delineates clear component‑level differences between acceleration and stability, and identifies architectural limits where omitting donor embeddings leads to significant performance loss.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.