arXiv Machine Learning

Few-Step Diffusion Language Models via Trajectory Self-Distillation

arXiv:2602. 12262v4 Announce Type: replace-cross Abstract: Diffusion large language models (DLLMs) have emerged as powerful generative models with the promise of fast text generation through parallel decoding.

arXiv Machine Learning
Sep 3

DLM-One: Diffusion Language Models for One-Step Sequence Generation

The paper introduces DLM-One, a score‑distillation framework that enables one‑step sequence generation with continuous diffusion language models (DLMs). By aligning a student model’s outputs with a pretrained teacher DLM’s score function in the forward‑diffused noisy space, DLM-One removes the need for iterative refinement. Experiments across various DLM architectures show up to ~2000× speedup in sampling steps and ~500× in wall‑clock time while retaining competitive performance, and the authors propose an adversarially‑regularized two‑stage training scheme to mitigate student degeneration.

By Tianqi Chen, Shujian Zhang, Mingyuan Zhou
arXiv AI
4d ago

Less Uniform Discrete Diffusion is More Powerful and Scalable

The paper introduces Less Uniform Diffusion (LUDI), a framework that improves uniform diffusion language models by using a less uniform loss and per-token time embeddings to guide reverse transitions and enable confidence-based few-step sampling. Experiments demonstrate that LUDI provides cleaner supervision, enhances few-step generation, and scales to a 7B model (LUDI-7B) that achieves a 3-token-per-step speedup over autoregressive decoding while matching masked diffusion baselines. The work suggests that UDLMs still have untapped potential for complex generation tasks.

By Kaibo Wang, Ding Ding, Fangyu Ding, Zijin Feng, Han Shi, Haili Bai, Jiacheng Sun, Yang Xiang
arXiv AI
Jun 15

Residual Context Diffusion Language Models

arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.

By Yuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi, Coleman Hooper, Jintao Zhang, Aditya Tomar, Michael W. Mahoney, Sewon Min, Mehrdad Farajtabar, Kurt Keutzer, Amir Gholami, Chenfeng Xu
arXiv Machine Learning
Jun 25

Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding

arXiv:2601. 17917v3 Announce Type: replace Abstract: Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirectional attention to achieve superior global coherence compared to autoregressive models.

By Zhongyu Xiao, Zhiwei Hao, Jianyuan Guo, Yong Luo, Jia Liu, Jie Xu, Han Hu
arXiv Computation and Language
Aug 28

Dependency-Aware Revocable Decoding for Efficient Diffusion Large Language Model Inference

The paper introduces Dependency-Aware Revocable Decoding (DARD), a training‑free framework for diffusion large language models that separates tokens into masked, candidate, and unmasked states. DARD verifies candidate tokens using a selective context that excludes less reliable tokens and adaptively regulates their influence on subsequent decoding. Experiments on 12 textual and multimodal benchmarks across three open‑source dLLMs show that DARD improves the speed‑quality Pareto frontier, achieving a 2.71× speedup and a 4.35‑point CIDEr gain over Saber on Flickr30K.

By Wooje Park, Insu Lee, Minyoung Noh, Jaeyun Jang, Sungmin Lee, Kyuhong Shim, Byonghyo Shim
arXiv AI
Aug 11

Hybrid Policy Distillation for LLMs

arXiv:2604. 20244v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined choices of divergence direction, optimization strategy, and data regime.

By Wenhong Zhu, Ruobing Xie, Rui Wang, Pengfei Liu