Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
arXiv:2605. 29548v2 Announce Type: replace Abstract: Larger models learn tasks smaller models do not.
arXiv:2606. 03990v1 Announce Type: new Abstract: We investigate whether neuron populations within neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss.
arXiv:2605. 29548v2 Announce Type: replace Abstract: Larger models learn tasks smaller models do not.
arXiv:2604. 18827v2 Announce Type: replace-cross Abstract: Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision.
arXiv:2606. 25008v1 Announce Type: new Abstract: Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute.
The paper investigates the often-overlooked scale vectors in large language models, showing that despite their tiny size they are crucial for pre‑training performance. The authors provide theoretical insights that scale vectors mainly aid optimization rather than expressivity, and they analyze how weight decay affects different normalization layers. Building on these findings, they propose lightweight improvements—branch‑specific heterogeneity, better placement, and magnitude‑direction reparameterization—that consistently reduce loss across a range of model sizes and training settings.
The paper studies Kolmogorov‑Arnold Networks (KANs), a neural architecture that treats activation functions as learnable components, offering improved interpretability for scientific applications. It investigates how KANs scale with dataset size on image classification tasks (MNIST, Fashion‑MNIST) and a magnetic‑parameter regression task, revealing a broken neural scaling law that transitions from a faster to a slower decay of test loss as data grows. The authors also analyze how the learned activation functions evolve from simple linear approximations to more complex, interpretable symbolic forms as more data is provided.
arXiv:2606. 29196v1 Announce Type: new Abstract: Do language models know when they are being tested?
arXiv:2509. 24882v2 Announce Type: replace Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models.
arXiv:2606. 07414v1 Announce Type: new Abstract: Sparsity allows scaling model parameters without proportionally increasing computational cost.
arXiv:2602. 07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset.
arXiv:2606. 04409v1 Announce Type: cross Abstract: Modern deep neural networks usually have large parameter scales and nonlinear hierarchical structures, and they have achieved strong performance in computer vision.
arXiv:2602. 15253v2 Announce Type: replace Abstract: Neural scaling laws -- power-law relationships between loss, model size, and data -- have been extensively documented for language and vision transformers, yet their existence in single-cell genomics remains largely unexplored.