arXiv:2605.06272v2 Announce Type: replace
Abstract: While generative modeling has achieved remarkable success on tasks like natural language-conditioned image generation, enabling model adaptation fr...
By Tyler Ingebrand, Ruihan Zhao, Kushagra Gupta, David Fridovich-Keil, Sandeep P. Chinchali, Ufuk Topcu
arXiv:2609.35763v3 Announce Type: replace
Abstract: Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representa...
By Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
arXiv:2505. 04486v4 Announce Type: replace-cross Abstract: Flow matching models have shown great potential in image generation tasks among probabilistic generative models.
By Anirban Samaddar, Yixuan Sun, Viktor Nilsson, Sandeep Madireddy
arXiv:2608.27885v1 Announce Type: new
Abstract: Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not dire...
By Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon
The paper introduces CARE, a lightweight, plug‑and‑play regularization framework for diffusion models that dynamically adjusts feature distributions based on condition similarity. By leveraging built‑in conditioning signals such as labels or text prompts, CARE promotes tighter feature clusters for similar conditions without requiring explicit alignment losses or external supervision. Empirical results show consistent improvements in visual fidelity and convergence stability, achieving significant FID reductions and speed‑ups on ImageNet and text‑to‑image tasks, and it can be combined with existing regularization methods for further gains.
By Fengjia Guo, Zhuoyi Yang, Jie Tang
The paper introduces a post‑training approach that enables a single inference process to transition from text reasoning to image synthesis, eliminating the need for explicit modality switching. Using the 14B BAGEL model, the authors demonstrate that targeted post‑training data and reward‑weighted training improve multimodal image generation across four independent T2I benchmarks. The study highlights the benefits of joint text‑image generation and strategic data selection for enhancing T2I performance.
By Jiahui Chen, Philippe Hansen-Estruch, Xiaochuang Han, Yushi Hu, Emily Dinan, Amita Kamath, Michal Drozdzal, Reyhane Askari-Hemmat, Luke Zettlemoyer, Marjan Ghazvininejad
Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions.
Diffusion and flow matching models have made significant progress in text-to-image generation, yet high computation, quadratic complexity, and large memory footprint hinder high-resolution synthesis and edge deployment. We propose Nexus, which integrates sparse architecture, linear complexity, and low-bit quantization.
The paper introduces dFlowGRPO, a reinforcement learning framework tailored for discrete flow models (DFMs). It generalizes previous work on diffusion large language models by supporting various probability paths and non-masked source distributions, and formulates denoising as a Markov decision process that leverages transition rates and posterior models. Experiments on the multimodal DFM FUDOKI show that dFlowGRPO outperforms existing GRPO methods on text‑to‑image generation and matches continuous flow models on multimodal understanding tasks.
By Zhengyan Wan, Yidong Ouyang, Panwen Hu, Qiang Sun
arXiv:2511.21415v2 Announce Type: replace
Abstract: We introduce DiverseVAR, a framework that enhances the diversity of text-conditioned visual autoregressive models (VAR) at test time without requir...
By Mingue Park, Prin Phunyaphibarn, Phillip Y. Lee, Minhyuk Sung
arXiv:2603. 12893v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learning from reward signals to explicitly improve desirable aspects such as image quality and prompt alignment.
By David McAllister, Miika Aittala, Tero Karras, Janne Hellsten, Angjoo Kanazawa, Timo Aila, Samuli Laine
The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
By Jaymin Bhan, JiHong Jeon, SangYeop Jeong