Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining
Read the original on arXiv AI →The article proposes treating large language model (LLM) data mixing as a classical mixture experiment, where data domains are components, token shares are proportions, and proxy-training runs serve as design points. Using sparse second‑order Scheffé response‑surface models, the authors construct model‑robust Σ‑optimal designs that efficiently identify optimal data mixtures and reveal strong interaction effects, especially between weak domains and web‑derived text. Empirical results on RegMix show that these designs recover mixture rankings while reducing proxy runs by about 25%, demonstrating that data mixing can be optimized through experimental design rather than solely prediction.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.