The article proposes treating large language model (LLM) data mixing as a classical mixture experiment, where data domains are components, token shares are proportions, and proxy-training runs serve as design points. Using sparse second‑order Scheffé response‑surface models, the authors construct model‑robust Σ‑optimal designs that efficiently identify optimal data mixtures and reveal strong interaction effects, especially between weak domains and web‑derived text. Empirical results on RegMix show that these designs recover mixture rankings while reducing proxy runs by about 25%, demonstrating that data mixing can be optimized through experimental design rather than solely prediction.
By Yicheng Mao, Hongru Du
arXiv:2602. 06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts.
By Junqi Chen, Sirui Chen, Chaochao Lu
arXiv:2603.26164v2 Announce Type: replace-cross
Abstract: Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters...
By Hao Liang, Zhengyang Zhao, Mingrui Chen, Meiyi Qiang, Lu Ma, Rongyi Yu, Hengyi Feng, Shixuan Sun, Zimo Meng, Xiaochen Ma, Xuanlin Yang, Qifeng Cai, Ruichuan An, Bohan Zeng, Zhen Hao Wong, Chengyu Shen, Runming He, Zhaoyang Han, Yaowei Zheng, Fangcheng Fu, Conghui He, Bin Cui, Zhiyu Li, Weinan E, Wentao Zhang
TabCausal is a causal discovery foundation model that learns to map datasets directly to causal graphs by pretraining across diverse causal environments. It uses a dynamic task construction strategy to expose the model to varied graph priors, mechanisms, noise models, dimensions, sample sizes, and intervention regimes, improving transferability from observational and mixed‑interventional data. On large synthetic benchmarks and a new protocol‑guided semantic benchmark, TabCausal outperforms many classical baselines and shows robust structure recovery, especially when interventional evidence is available.
By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv:2607. 22769v1 Announce Type: cross Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data.
By He Zhang
arXiv:2607. 11052v1 Announce Type: new Abstract: Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential.
By Kimia Hamidieh, Lester Mackey, David Alvarez-Melis