ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling
arXiv:2607. 19332v1 Announce Type: new Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching.
Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold.
arXiv:2607. 19332v1 Announce Type: new Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching.
arXiv:2607. 27372v1 Announce Type: new Abstract: The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages.
arXiv:2509. 24935v3 Announce Type: replace-cross Abstract: Scalability has driven recent advances in generative modeling, yet its principles remain underexplored for adversarial learning.
arXiv:2609.36348v1 Announce Type: cross Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
arXiv:2605. 00941v4 Announce Type: replace Abstract: Flow matching has become a leading framework for generative modeling, but quantifying the uncertainty of its samples remains an open problem.
LLaDA-Image is a unified framework that couples a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision‑language module based on the LLaDA2.0‑Mini diffusion language model. The approach first builds a strong visual generative prior through image‑only pre‑training and mid‑training, then fine‑tunes with a 220M‑sample generation pipeline that includes 98 real images. The resulting model produces highly photorealistic images that accurately follow fine‑grained editing instructions, and a distilled version, LLaDA‑Image‑Turbo, enables fast inference in 2–4 sampling steps. On Qwen‑Image‑Bench, LLaDA‑Image sets new state‑of‑the‑art scores for open‑source models in both English and Chinese tracks, and the authors release weights, code, and detailed recipes to support further research.
arXiv:2609.38920v1 Announce Type: new Abstract: Offline multi-objective optimization (MOO) seeks solutions with better objective trade-offs using only a fixed dataset, without querying the objectives...
arXiv:2605. 29920v2 Announce Type: replace Abstract: We introduce Midpoint Generative Models (MGM), a principled framework for training one-step generative models.
arXiv:2609.37147v1 Announce Type: cross Abstract: Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule...
The paper investigates training objectives for denoising-based generative models, focusing on loss weighting and output parameterization such as noise-, clean image-, and velocity-based formulations. It conducts a systematic numerical study across synthetic datasets with controlled geometry and real image data, evaluating denoising accuracy via PSNR and generative quality via FID. The goal is to disentangle how training choices interact with data manifold dimensionality, model architecture, and dataset size, offering practical design insights rather than proposing a new method.
arXiv:2602.19600v2 Announce Type: replace Abstract: Many high-dimensional datasets concentrate near a low-dimensional structure embedded in the ambient space. Generative models for such data must con...
The paper introduces continuous adversarial flow models, a continuous-time flow framework trained with an adversarial objective that replaces the fixed mean-squared-error criterion of flow matching. By incorporating a learned discriminator, the method guides training toward a different generalized distribution, yielding samples more closely aligned with the target data distribution. Applied as a post‑training step, it markedly improves ImageNet 256px generation metrics—reducing the guidance‑free FID of latent‑space SiT from 8.26 to 3.63 and of pixel‑space JiT from 7.17 to 3.57—and also enhances guided generation and text‑to‑image benchmarks.