arXiv Machine Learning By Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi

Small-Scale Experiments: Are We There Yet?

Read the original on arXiv Machine Learning →

arXiv:2608. 11859v1 Announce Type: new Abstract: Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 12

Small-Scale Experiments: Are We There Yet?

Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided.

arXiv Machine Learning
Sep 23

Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World

The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.

By Christopher M. Bryant, Hao Liu
arXiv Machine Learning
1d ago

Learning Rate Transfer for Hybrid Transformer-SSM Architectures

arXiv:2610. 01172v1 Announce Type: new Abstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models.

By Jimin Seo, Gyubok Lee, Yeonsik Jo, Kiwoong Yoo, Yeongoon Kim, Minhae Oh, Jin Woo Koo, Suhwan Kim, Nakyung Lee, Minsik Seol, Idris Nechnech, Jaehyeon Kim, Giho Lee, Jungwoo Lee