DiffusionGemma: 4x faster text generation
Related stories
Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
arXiv:2603. 12506v2 Announce Type: replace-cross Abstract: Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise.
Benchmarking Text Generation Inference
Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models
The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.
Introducing Würstchen: Fast Diffusion for Image Generation
Just on Time: Token-Level Early Stopping for Diffusion Language Models
arXiv:2602. 11133v2 Announce Type: replace Abstract: Diffusion language models generate text through iterative refinement, a process that is often computationally inefficient because many tokens reach stability long before the final denoising step.
A Survey on Diffusion Language Models
arXiv:2508. 10875v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm.
Faster Text Generation with Self-Speculative Decoding
Representation-based Masked Diffusion Model
The paper introduces Representation-based Masked Diffusion Model (RMDM), a new framework for language modeling that improves upon existing Masked Diffusion Models by incorporating global semantic guidance. RMDM encodes text into a continuous semantic space with a pretrained encoder, normalizes this representation to a Gaussian prior via an invertible transformation, and then trains a masked diffusion model conditioned on this latent representation to coordinate parallel token updates. Experiments show that RMDM yields higher generation quality, especially when using aggressive few‑step sampling.
Welcome aMUSEd: Efficient Text-to-Image Generation
DICE: Distilling Classifier-Free Guidance into Text Embeddings
arXiv:2502.03726v3 Announce Type: replace Abstract: Text-to-image diffusion models are capable of generating high-quality images, but suboptimal pre-trained text representations often result in these...