Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
arXiv:2511. 09002v3 Announce Type: replace-cross Abstract: Self-consuming generative models have received significant attention over the last few years.
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
arXiv:2601. 21868v2 Announce Type: replace-cross Abstract: Understanding the stability and long-time behavior of generative models is a fundamental problem in modern machine learning.
arXiv:2502. 18049v5 Announce Type: replace-cross Abstract: Recent studies identified an intriguing phenomenon in recursive generative model training known as model collapse, where models trained on data generated by previous models exhibit severe performance degradation.
arXiv:2610.01318v1 Announce Type: cross Abstract: Model collapse arises when generative models are trained on synthetic data produced by earlier models. The phenomenon has attracted considerable atte...
arXiv:2606. 08953v1 Announce Type: new Abstract: Modern generative models often define an entire probability path from a simple prior to the data law, rather than only an endpoint map.
arXiv:2608. 12438v1 Announce Type: new Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action.
arXiv:2605.29713v2 Announce Type: replace-cross Abstract: This book provides a compact, derivation-oriented introduction to the mathematical foundations of modern generative artificial intelligence....
arXiv:2609.15193v1 Announce Type: new Abstract: Drifting models offer a promising route to faster generative AI: they perform gradual transport during training, while generating new samples in a sing...
arXiv:2605. 07724v2 Announce Type: replace-cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal, the model tends to collapse onto a narrow set of outputs that over-optimize that objective.
The paper introduces Geometrically Modified Outputs (GMOs), a technique that reweights the singular values of a generative model’s input-output Jacobian to strengthen the negative signal used in self‑training. By amplifying mode‑seeking behavior and distortions in standard outputs, GMOs provide a more targeted negative guidance for models such as Neon and SIMS. Experiments on one‑step generative models show that GMOs consistently improve the performance of negative‑guidance self‑training methods compared to using unmodified outputs.
arXiv:2607. 15623v1 Announce Type: cross Abstract: Predictive models deployed at scale influence future data, a phenomenon called performativity.