Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
arXiv:2310. 05264v5 Announce Type: replace Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility": given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs.
arXiv:2608.21366v1 Announce Type: new Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors....
arXiv:2502. 18049v5 Announce Type: replace-cross Abstract: Recent studies identified an intriguing phenomenon in recursive generative model training known as model collapse, where models trained on data generated by previous models exhibit severe performance degradation.
arXiv:2608. 02575v1 Announce Type: new Abstract: Diffusion models rely on stochastic inputs, yet on finite-precision hardware, the "randomness" they consume is realized as deterministic numerical orbits generated by pseudorandom rules.
The paper introduces Abra, a family of flow‑matching transformers used to systematically study scaling laws for text‑to‑image diffusion models across three orders of magnitude in compute. It finds that diffusion models scale predictably like language models but need far more data, with compute optimality occurring at roughly 200 image tokens per parameter—ten times the optimal ratio for large language models. The study also shows that diffusion models are robust to overtraining, that more data is preferable to larger models, and that scaling predictability extends to generative quality, optimal CFG settings, representation quality, and training curve shapes.
arXiv:2607. 20913v1 Announce Type: new Abstract: Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained model's broad generative capability.
arXiv:2510. 17917v2 Announce Type: replace-cross Abstract: Data unlearning aims to remove the influence of specific training samples from a trained model.
arXiv:2609.39124v1 Announce Type: new Abstract: Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many spe...
arXiv:2507. 04219v5 Announce Type: replace-cross Abstract: Current unlearning methods for LLMs optimize on the private information they seek to remove by incorporating it into their fine-tuning data.
arXiv:2606. 23920v1 Announce Type: cross Abstract: The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to produce samples from compositionally-defined target distributions such as a geometric combination of the source distributions.
arXiv:2609.37974v1 Announce Type: cross Abstract: Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The m...