arXiv AI By Yuxin Ma, Nan Chen, Mateo D\'iaz, Soufiane Hayou, Dmitriy Kunisky, Soledad Villar

$\mu$pscaling small models: Principled warm starts and hyperparameter transfer

Read the original on arXiv AI →

arXiv:2602. 10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Statistics ML
2d ago

Transferable Graph Metanetworks

arXiv:2610.00420v1 Announce Type: new Abstract: A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such...

By Yuxin Ma, Adir Dayan, Yam Eitan, Haggai Maron, Soledad Villar
Hugging Face Trending Papers
Aug 12

Small-Scale Experiments: Are We There Yet?

Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided.