arXiv Machine Learning

Thermodynamic Weight Decay: Exploring Grokking Acceleration via Attention Specific Heat

arXiv:2607. 20552v1 Announce Type: new Abstract: Grokking -- the delayed generalization of neural networks long after they have memorized their training data -- wastes thousands of training epochs and is notoriously unpredictable.

Hugging Face Trending Papers
Aug 4

Predicting Deep Neural Network Training Outcomes from Early Training Telemetry

Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single training run's own early telemetry - per-epoch loss, training accuracy, gradient signal-to-noise ratio, weight-norm growth, and an activation-saturation snapshot - together with its sampled hyperparameters, can predict that run's eventual outcome without reference to other runs.