CausalMix: Data Mixture as Causal Inference for Language Model Training
arXiv:2607. 01104v1 Announce Type: cross Abstract: In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance.
arXiv:2606. 04401v1 Announce Type: new Abstract: The capabilities of large language models (LLMs) significantly depend on training data drawn from various domains.
arXiv:2607. 01104v1 Announce Type: cross Abstract: In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance.
arXiv:2607. 11052v1 Announce Type: new Abstract: Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential.
The article proposes treating large language model (LLM) data mixing as a classical mixture experiment, where data domains are components, token shares are proportions, and proxy-training runs serve as design points. Using sparse second‑order Scheffé response‑surface models, the authors construct model‑robust Σ‑optimal designs that efficiently identify optimal data mixtures and reveal strong interaction effects, especially between weak domains and web‑derived text. Empirical results on RegMix show that these designs recover mixture rankings while reducing proxy runs by about 25%, demonstrating that data mixing can be optimized through experimental design rather than solely prediction.
arXiv:2607. 22769v1 Announce Type: cross Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data.
arXiv:2604. 26170v2 Announce Type: replace Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge.
arXiv:2604. 03532v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult.
arXiv:2606. 16456v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models enable efficient scaling, but training them from scratch remains prohibitively expensive.
arXiv:2606. 14971v1 Announce Type: cross Abstract: While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture for pre-training and post-training remains a significant open problem.
arXiv:2607. 24516v1 Announce Type: cross Abstract: While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed.
arXiv:2607. 27928v1 Announce Type: new Abstract: The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data.
arXiv:2608. 14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)).
arXiv:2607. 22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes.