arXiv:2606. 18694v1 Announce Type: new Abstract: A network of oscillators that synchronizes perfectly computes nothing further, so an attention architecture built from synchronization must locate its computation in structured departures from agreement.
By Joshua Nunley
arXiv:2607. 05104v1 Announce Type: cross Abstract: Grokking -- the delayed onset of generalization long after a network has fit its training set - -is usually studied in models too large to read completely and reported from single training runs.
By Yoshiyuki Ootani
The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.
By Anish Kataria
arXiv:2607. 06639v1 Announce Type: cross Abstract: On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized.
By Truong Xuan Khanh
arXiv:2607. 04333v1 Announce Type: new Abstract: Grokking -- generalization arriving long after training-set interpolation -- can be accelerated by structure-agnostic interventions: gradient filtering, weight-norm clamping, geometric penalties on hidden representations.
By Gunner Levi Howe
The paper demonstrates that emergent capabilities in machine learning models can be forecasted with lead time, calibrated uncertainty, and controlled false‑alarm rates. Using per‑seed analysis on transformers, the authors show that the formation time of a previous‑token head predicts the emergence of an induction head with Spearman ρ = 0.977 and a median lead of 975 training steps. Conformal intervals, blind pre‑registered tests, and a multiplicative rule relating anchor and event times further validate the predictive framework across multiple model families and configurations.
By Gunner Levi Howe