The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.
By Christopher M. Bryant, Hao Liu
arXiv:2609.37745v1 Announce Type: new
Abstract: Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent...
By Hyunseok Lee, Mihir Basil, Yizhou Liu, Jeff Gore
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2607. 14144v2 Announce Type: replace Abstract: The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality.
By Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong
arXiv:2608. 10288v1 Announce Type: new Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws.
By Burc Gokden
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie