arXiv Machine Learning By Yizhou Liu, Jeff Gore

Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

Read the original on arXiv Machine Learning →

arXiv:2606. 25008v1 Announce Type: new Abstract: Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 7

Deriving Neural Scaling Laws from the statistics of natural language

arXiv:2602. 07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset.

By Francesco Cagnetta, Allan Ravent\'os, Surya Ganguli, Matthieu Wyart