arXiv Machine Learning By Matteo Marchi, Jo\~ao Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada

Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data

Read the original on arXiv Machine Learning →

The paper investigates how to avoid model collapse when training large language models with synthetic data. It establishes theoretical guarantees for the minimum ratio of human to synthetic data needed to maintain training stability, using the Fisher‑Rao metric to analyze dynamics on the probability simplex. The authors derive contraction and invariance bounds that remain meaningful even in high dimensions, showing that the required data ratio differs from earlier estimates.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.