arXiv Machine Learning By Alex Buna, Fanghui Liu, Patrick Rebeschini

Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent

Read the original on arXiv Machine Learning →

The paper studies how to allocate a fixed computational budget between pretraining and fine‑tuning in a two‑stage ridge regression setting. By modeling the process as a compute‑split problem and analyzing data‑dependent evaluation geometries, it derives the optimal split in terms of prediction‑relevant spectral components of the pretraining and fine‑tuning empirical covariances. The analysis uses a basis‑invariant eigenspace decomposition and perturbative control of non‑commuting dynamics to capture how pretraining directions influence downstream predictions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 7

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

arXiv:2607. 04033v1 Announce Type: cross Abstract: Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the landscape of over one hundred methods remains fragmented.

By Siyuan Li, Jiabao Pan, Yumou Liu, Zhuoli Ouyang, Xin Jin, Xinglong Xu, Jingxuan Wei, Shengye Pang, Jintao Che, Xuanhe Zhou, Conghui He, Cheng Tan