Small-Scale Experiments: Are We There Yet?
arXiv:2608. 11859v1 Announce Type: new Abstract: Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver.
Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided.
arXiv:2608. 11859v1 Announce Type: new Abstract: Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver.
arXiv:2602. 10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets.
The paper proposes a framework to reduce the cost of constructing scaling laws for large foundation models by treating data collection as a Bayesian optimization problem. It shows that expanding the compute budget progressively and augmenting observed configurations with surrogate-fantasized evaluations can recover a broad experimental grid, enabling accurate scaling law fitting without training every configuration. This approach can achieve computational savings of up to 10–100× compared to a full dense grid.
The paper introduces Power‑Law Entropy Search (PLES), a computational‑cost‑aware acquisition function that uses multi‑fidelity Bayesian optimization to efficiently estimate optimal hyperparameter scaling laws for large language model training. PLES focuses on reducing the overall uncertainty of scaling law estimates rather than optimizing a single objective, selecting configurations that maximize uncertainty reduction per unit computational cost. Experiments on synthetic benchmarks, surrogate models, and real LLM pre‑training runs show that PLES converges to accurate scaling laws using less than one‑tenth of the computational budget required by conventional grid search and other baselines.
The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.
The paper investigates the often-overlooked scale vectors in large language models, showing that despite their tiny size they are crucial for pre‑training performance. The authors provide theoretical insights that scale vectors mainly aid optimization rather than expressivity, and they analyze how weight decay affects different normalization layers. Building on these findings, they propose lightweight improvements—branch‑specific heterogeneity, better placement, and magnitude‑direction reparameterization—that consistently reduce loss across a range of model sizes and training settings.
arXiv:2608. 20061v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost.
arXiv:2606. 07616v1 Announce Type: cross Abstract: Scaling laws provide a fundamental framework for understanding the performance of Language Models (LMs), yet deriving them requires prohibitively expensive evaluations across thousands of checkpoints or millions of inference samples.
arXiv:2609.08690v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyp...
arXiv:2604. 22753v2 Announce Type: replace Abstract: Scaling laws are used to plan multi-million-dollar training runs, but fitting those laws can itself cost millions.
arXiv:2605. 29548v2 Announce Type: replace Abstract: Larger models learn tasks smaller models do not.
arXiv:2610. 01172v1 Announce Type: new Abstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models.