arXiv:2502. 18049v5 Announce Type: replace-cross Abstract: Recent studies identified an intriguing phenomenon in recursive generative model training known as model collapse, where models trained on data generated by previous models exhibit severe performance degradation.
By Hengzhi He, Shirong Xu, Guang Cheng
The paper introduces SynPro, a synthetic data generation framework that augments limited organic text for large language model pretraining by applying rephrasing and reformatting operations. SynPro’s generators are optimized with reinforcement learning rewards for quality, faithfulness, and data influence, and are updated continuously as training plateaus. Experiments on 400M, 1.1B, and 2B models show that SynPro can unlock 3.4–5.2× the effective tokens of standard repetition, even outperforming a non‑data‑bound oracle at larger scales.
By Zichun Yu, Chenyan Xiong
arXiv:2609.09572v1 Announce Type: new
Abstract: Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Doh...
By Jichu li, Difan Zou
arXiv:2507. 04219v5 Announce Type: replace-cross Abstract: Current unlearning methods for LLMs optimize on the private information they seek to remove by incorporating it into their fine-tuning data.
By Yan Scholten, Sophie Xhonneux, Leo Schwinn, Stephan G\"unnemann
arXiv:2608.21366v1 Announce Type: new
Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors....
By Xihao Xie, Beichen Hu
arXiv:2605. 12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety.
By Vin\'icius Conte Turani, Ot\'avio Parraga, Jo\~ao Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinsk\"u
arXiv:2607. 22994v1 Announce Type: cross Abstract: Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting.
By Tao Zhang, Qixuan Fan, Yiyuan Liang, Yanjie Wang, Song Yan, Tian Tian, Jiahuan Zhou, Luxin Yan, Sheng Zhong, Xu Zou
arXiv:2610.01318v1 Announce Type: cross
Abstract: Model collapse arises when generative models are trained on synthetic data produced by earlier models. The phenomenon has attracted considerable atte...
By Hanna Malet, Gabriel Turinici
The paper investigates how to avoid model collapse when training large language models with synthetic data. It establishes theoretical guarantees for the minimum ratio of human to synthetic data needed to maintain training stability, using the Fisher‑Rao metric to analyze dynamics on the probability simplex. The authors derive contraction and invariance bounds that remain meaningful even in high dimensions, showing that the required data ratio differs from earlier estimates.
By Matteo Marchi, Jo\~ao Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada
arXiv:2410. 12341v4 Announce Type: replace-cross Abstract: As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a process known as AI autophagy.
By Daniele Gambetta, Gizem Gezici, Fosca Giannotti, Dino Pedreschi, Alistair Knott, Luca Pappalardo
arXiv:2607. 15623v1 Announce Type: cross Abstract: Predictive models deployed at scale influence future data, a phenomenon called performativity.
By Moritz Hardt
arXiv:2603. 11784v2 Announce Type: replace Abstract: As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publicly available online text may be consumed.
By Giorgio Racca, Michal Valko, Amartya Sanyal