The paper introduces DLM-One, a score‑distillation framework that enables one‑step sequence generation with continuous diffusion language models (DLMs). By aligning a student model’s outputs with a pretrained teacher DLM’s score function in the forward‑diffused noisy space, DLM-One removes the need for iterative refinement. Experiments across various DLM architectures show up to ~2000× speedup in sampling steps and ~500× in wall‑clock time while retaining competitive performance, and the authors propose an adversarially‑regularized two‑stage training scheme to mitigate student degeneration.
By Tianqi Chen, Shujian Zhang, Mingyuan Zhou
arXiv:2602. 19066v2 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) have recently achieved strong results in text generation.
By David Li, Nikita Gushchin, Dmitry Abulkhanov, Eric Moulines, Ivan Oseledets, Maxim Panov, Alexander Korotin
The paper introduces a new method for training data attribution in diffusion models called TID, which uses a local score discrepancy measure and can be estimated without retraining. It further distills this approach into TIDE, a forward‑only student that reproduces the teacher’s rankings using internal activations, achieving comparable accuracy at dramatically lower query cost. Experiments on CIFAR‑10, ArtBench‑10, and MS‑COCO show that TID outperforms existing methods and TIDE attributes samples in milliseconds, faster than generation itself.
By Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji
arXiv:2606. 29869v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution fitting with long-tail probability modeling, limiting both generation quality and generalization.
By Zilong Liu, Xuewen Zhang, Jinrui Xing, Juyi Qiao, Huiyong Wang, Junming Jiao
arXiv:2603. 22216v2 Announce Type: replace-cross Abstract: The slow, sequential nature of autoregressive (AR) language models has driven the adoption of parallel decoding methods.
By Chi Zhang, Xixi Hu, Bo Liu, Qiang Liu
arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.
By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv:2606. 09456v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models.
By Yifan Niu, Han Xiao, Dongyi Liu, Zelong Wang, Dihong Gong, Yasheng Wang, Jia Li
arXiv:2609. 04010v1 Announce Type: new Abstract: Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation.
By Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu
The paper revisits the continuous diffusion language model Plaid and introduces RePlaid, aligning its architecture with modern discrete diffusion models. RePlaid achieves a compute gap of only 20× compared to autoregressive models, surpasses Duo with fewer parameters, and outperforms MDLM in over‑trained settings. On OpenWebText, RePlaid sets a new state‑of‑the‑art continuous diffusion perplexity of 22.1 and demonstrates superior generation quality, while theoretical analysis links likelihood‑based training to linear cross‑entropy over time and structured embedding geometries.
By Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sekhar Sahoo, Yongxin Chen, Arash Vahdat, Morteza Mardani, John Thickstun
arXiv:2607. 21427v1 Announce Type: new Abstract: Discrete flow matching provides a flexible framework for generative modeling on discrete structures.
By Daniil Cherniavskii, Daniel Severo, Karen Ullrich
arXiv:2607. 19686v1 Announce Type: cross Abstract: Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging.
By Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, Lexing Ying
arXiv:2609.38364v1 Announce Type: cross
Abstract: Discrete diffusion models and flow matching have emerged as powerful frameworks for generative modeling over discrete state spaces, yet efficient few...
By Yidong Ouyang, Zhengyan Wan, Themis Haris, Tian Tan, Liqian Peng, Henry Li, Ziqian Lin, Jianhang Chen, Maryam Karimzadehgan, Alec Go, George Michailidis