ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling
arXiv:2607. 19332v1 Announce Type: new Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching.
arXiv:2607. 19332v1 Announce Type: new Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching.
arXiv:2609.37080v1 Announce Type: new Abstract: Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusi...
arXiv:2609.30988v1 Announce Type: new Abstract: Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while...
Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold.
Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While methods improve the FID and related metrics, it is increasingly unclear whether they reflect real progress in generative modeling.
arXiv:2512. 19311v2 Announce Type: replace-cross Abstract: This paper studies the training-testing discrepancy (a.
arXiv:2502. 10389v2 Announce Type: replace-cross Abstract: Diffusion models (DMs) have become the leading choice for generative tasks across diverse domains.
arXiv:2604. 26985v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) generate discrete sequences by iterative denoising under an absorbing masking process.
LLaDA-Image is a unified framework that couples a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision‑language module based on the LLaDA2.0‑Mini diffusion language model. The approach first builds a strong visual generative prior through image‑only pre‑training and mid‑training, then fine‑tunes with a 220M‑sample generation pipeline that includes 98 real images. The resulting model produces highly photorealistic images that accurately follow fine‑grained editing instructions, and a distilled version, LLaDA‑Image‑Turbo, enables fast inference in 2–4 sampling steps. On Qwen‑Image‑Bench, LLaDA‑Image sets new state‑of‑the‑art scores for open‑source models in both English and Chinese tracks, and the authors release weights, code, and detailed recipes to support further research.
The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.
arXiv:2506. 14753v3 Announce Type: replace-cross Abstract: Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process.
arXiv:2606. 00094v1 Announce Type: cross Abstract: Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, low-dimensional, and compact parameterization space.