The paper investigates how to avoid model collapse when training large language models with synthetic data. It establishes theoretical guarantees for the minimum ratio of human to synthetic data needed to maintain training stability, using the Fisher‑Rao metric to analyze dynamics on the probability simplex. The authors derive contraction and invariance bounds that remain meaningful even in high dimensions, showing that the required data ratio differs from earlier estimates.
By Matteo Marchi, Jo\~ao Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
By Bingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng Xu
arXiv:2609.09572v1 Announce Type: new
Abstract: Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Doh...
By Jichu li, Difan Zou
arXiv:2610.00873v1 Announce Type: cross
Abstract: In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between...
By Hongyu Cao, Xinyuan Wang, Arun Vignesh Malarkkan, Kunpeng Liu, Haifeng Chen, Yanjie Fu
arXiv:2606. 13796v1 Announce Type: cross Abstract: Recursive training of generative models on their own outputs can lead to model collapse, a compounding drift away from the true data distribution.
By Na\"il B. Khelifa, Richard E. Turner, Ramji Venkataramanan
arXiv:2606. 15959v1 Announce Type: cross Abstract: Neural networks are used as generative surrogate models for scientific discovery, which are trainable approximations of scientific simulations.
By Zhimin Li, Harshitha Menon, Charles Jekel, Valerio Pascucci, Peter Lindstrom