arXiv AI By Kevin Wang, Hongqian Niu, Didong Li

Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training

Read the original on arXiv AI →

arXiv:2602. 16065v2 Announce Type: replace-cross Abstract: As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive degradation or collapse.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data

The paper investigates how to avoid model collapse when training large language models with synthetic data. It establishes theoretical guarantees for the minimum ratio of human to synthetic data needed to maintain training stability, using the Fisher‑Rao metric to analyze dynamics on the probability simplex. The authors derive contraction and invariance bounds that remain meaningful even in high dimensions, showing that the required data ratio differs from earlier estimates.

By Matteo Marchi, Jo\~ao Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada