Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
The paper introduces Geometrically Modified Outputs (GMOs), a technique that reweights the singular values of a generative model’s input-output Jacobian to strengthen the negative signal used in self‑training. By amplifying mode‑seeking behavior and distortions in standard outputs, GMOs provide a more targeted negative guidance for models such as Neon and SIMS. Experiments on one‑step generative models show that GMOs consistently improve the performance of negative‑guidance self‑training methods compared to using unmodified outputs.
arXiv:2510. 16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models.
arXiv:2607. 27372v1 Announce Type: new Abstract: The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages.
arXiv:2512.17303v3 Announce Type: replace Abstract: In diffusion and flow-matching generative models, guidance techniques are widely used to improve sample quality and consistency. Classifier-free gu...
arXiv:2608. 07924v1 Announce Type: cross Abstract: Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined sample-based drift field.
arXiv:2510. 17136v2 Announce Type: replace Abstract: The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models.
arXiv:2606. 08578v1 Announce Type: new Abstract: Recently, large time series models (LTSMs) have gained increasing attention due to their similarities to large language models, including flexible context length, scalability, and task generality, outperforming advanced task-specific models.
SelfLift is a progressive‑resolution framework that accelerates few‑step diffusion models by enabling late, self‑recovering transitions between low‑ and high‑resolution latents. It introduces a training‑free Artifact‑Aware Consistency Lift that uses disagreement between direct latent lifting and pixel‑VAE re‑encoding to detect and correct artifacts, and a self‑recovery policy that transfers high‑resolution guidance from an internal teacher. Experiments on FLUX.2‑Klein and Z‑Image‑Turbo show latency reductions of 41.5% and 44.1%, and overall speedups of 29.61× and 19.21× over 50‑step baselines while maintaining competitive generation quality.
arXiv:2608. 13932v1 Announce Type: new Abstract: Iterative Generative Models (IGMs) span autoregressive and diffusion paradigms, and hybrid variants that couple them can achieve remarkable image-generation fidelity.
We’ve made progress towards stable and scalable training of energy-based models (EBMs) resulting in better sample quality and generalization ability than existing models. Generation in EBMs spends more compute to continually refine its answers and doing so can generate samples competitive with GANs at low temperatures, while also having mode coverage guarantees of likelihood-based models.
arXiv:2605. 29547v2 Announce Type: replace-cross Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators.
arXiv:2606. 05264v1 Announce Type: new Abstract: Training robust multivariate time series forecasting models requires large, diverse corpora, yet many real-world domains provide only a handful of observed sequences.
arXiv:2410. 02596v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found applications in various important tasks, attracting great research interest in their training algorithms.