arXiv Machine Learning

ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

arXiv:2607. 19332v1 Announce Type: new Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching.

arXiv Machine Learning
4d ago

Improved Distributional Diffusion Models

arXiv:2609.37147v1 Announce Type: cross Abstract: Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule...

By Tommaso Martorella, Alexandre Galashov, Felix Krause, Stefan Andreas Baumann, Valentin De Bortoli, Arthur Gretton, Bj\"orn Ommer
arXiv AI
Sep 4

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

LLaDA-Image is a unified framework that couples a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision‑language module based on the LLaDA2.0‑Mini diffusion language model. The approach first builds a strong visual generative prior through image‑only pre‑training and mid‑training, then fine‑tunes with a 220M‑sample generation pipeline that includes 98 real images. The resulting model produces highly photorealistic images that accurately follow fine‑grained editing instructions, and a distilled version, LLaDA‑Image‑Turbo, enables fast inference in 2–4 sampling steps. On Qwen‑Image‑Bench, LLaDA‑Image sets new state‑of‑the‑art scores for open‑source models in both English and Chinese tracks, and the authors release weights, code, and detailed recipes to support further research.

By Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie
arXiv Computer Vision
Sep 18

Training Flow Matching: The Role of Weighting and Parameterization

The paper investigates training objectives for denoising-based generative models, focusing on loss weighting and output parameterization such as noise-, clean image-, and velocity-based formulations. It conducts a systematic numerical study across synthetic datasets with controlled geometry and real image data, evaluating denoising accuracy via PSNR and generative quality via FID. The goal is to disentangle how training choices interact with data manifold dimensionality, model architecture, and dataset size, offering practical design insights rather than proposing a new method.

By Anne Gagneux, S\'egol\`ene Martin, R\'emi Gribonval, Mathurin Massias
arXiv Machine Learning
Sep 17

Beyond Random Couplings: Contrastive Noise Alignment in Generative Flows

The paper introduces Contrastive Noise Alignment (CNA), a training-time method for generative flow models that dynamically aligns Gaussian noise with data samples using a cross-modal InfoNCE objective. By modeling noise as an interacting particle system and regularizing with angular entropy and radial norm penalties, CNA reduces arbitrary data-noise couplings and flow curvature. Empirical results show that CNA improves generation quality, lowering FID by over 50% for few-step pixel-space generation compared to standard rectified flow and outperforming optimal transport baselines by at least 24%.

By Lennart Wittke, Vinicius Azevedo