arXiv AI

$\mu$pscaling small models: Principled warm starts and hyperparameter transfer

arXiv:2602. 10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets.

arXiv Statistics ML
2d ago

Transferable Graph Metanetworks

arXiv:2610.00420v1 Announce Type: new Abstract: A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such...

By Yuxin Ma, Adir Dayan, Yam Eitan, Haggai Maron, Soledad Villar
Hugging Face Trending Papers
Aug 12

Small-Scale Experiments: Are We There Yet?

Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided.

arXiv AI
Sep 2

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

The paper introduces Power‑Law Entropy Search (PLES), a computational‑cost‑aware acquisition function that uses multi‑fidelity Bayesian optimization to efficiently estimate optimal hyperparameter scaling laws for large language model training. PLES focuses on reducing the overall uncertainty of scaling law estimates rather than optimizing a single objective, selecting configurations that maximize uncertainty reduction per unit computational cost. Experiments on synthetic benchmarks, surrogate models, and real LLM pre‑training runs show that PLES converges to accurate scaling laws using less than one‑tenth of the computational budget required by conventional grid search and other baselines.

By Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin
arXiv AI
Jul 3

Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

arXiv:2602. 03001v2 Announce Type: replace-cross Abstract: To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and costly to tune.

By Hiroki Naganuma, Shagun Gupta, Youssef Briki, Ioannis Mitliagkas, Irina Rish, Parameswaran Raman, Hao-Jun Michael Shi