arXiv Machine Learning

Recursive Scaling in Masked Diffusion Models

arXiv:2606. 18022v1 Announce Type: new Abstract: Masked diffusion models (MDMs) have recently emerged as a promising paradigm for sequence generation.

arXiv AI
Jul 20

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

arXiv:2511. 15927v4 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead.

By Vaibhav Singh, Oleksiy Ostapenko, Pierre-Andr\'e No\"el, Eugene Belilovsky, Torsten Scholak
arXiv Machine Learning
Aug 6

Esoteric Language Models: A Family of Any-Order Diffusion LLMs

arXiv:2506. 01928v5 Announce Type: replace-cross Abstract: Diffusion-based language models offer a compelling alternative to autoregressive (AR) models by enabling parallel and controllable generation.

By Subham Sekhar Sahoo, Zhihan Yang, Yash Akhauri, Johnna Liu, Deepansha Singh, Zhoujun Cheng, Zhengzhong Liu, Eric Xing, John Thickstun, Arash Vahdat
arXiv AI
Jun 15

Residual Context Diffusion Language Models

arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.

By Yuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi, Coleman Hooper, Jintao Zhang, Aditya Tomar, Michael W. Mahoney, Sewon Min, Mehrdad Farajtabar, Kurt Keutzer, Amir Gholami, Chenfeng Xu
arXiv Computation and Language
Aug 28

Forward-Free Diffusion Language Models with BPTT-Free Looped Refinement

Forward-Free Diffusion Language Models with BPTT-Free Looped Refinement (FReDA) removes the need for a hand‑designed forward process in diffusion language modeling by treating model‑generated drafts as implicit intermediate states and refining them recursively. The approach detaches earlier refinement passes, backpropagating only through the final pass, and supports both self‑refinement and Best‑of‑N candidate selection. In sub‑8B experiments, FReDA‑4B surpasses larger diffusion baselines on reasoning and coding tasks, achieving up to 15% absolute gains and a 1.5‑1.8× speedup while scaling well with additional refinement steps.

By Haotian Sun, Rushi Qiang, Yuqian Zheng, Bo Dai
arXiv Machine Learning
Sep 2

Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture

The paper introduces Any-Order GPT, a Masked Diffusion Model (MDM) that operates within a decoder‑only architecture, allowing a fair comparison with traditional autoregressive (AR) models. By decoupling algorithmic choices from architectural differences, the study demonstrates that decoder‑only MDMs can achieve up to 25× faster inference while maintaining comparable perplexity, especially when using temperature annealing. The work offers insights into building more computationally efficient foundation models by separating core modeling decisions from architectural influences.

By Shuchen Xue, Tianyu Xie, Tianyang Hu, Zijin Feng, Jiacheng Sun, Kenji Kawaguchi, Zhenguo Li, Zhi-Ming Ma
arXiv Machine Learning
Aug 31

Learned Relay Representations for Forward-Thinking Discrete Diffusion Models

The paper introduces Learned Relay Representations (Relay), a technique for Masked Diffusion Models (MDMs) that preserves and forwards internal latent information across denoising steps via a differentiable per-token channel trained with truncated backpropagation through time. Relay enables MDMs to be forward‑thinking, avoiding costly recomputation of internal representations and improving efficiency. Applied to Fast‑dLLM v2, Relay outperforms standard supervised fine‑tuning on coding tasks and reduces inference latency by up to 32%, demonstrating a clear performance‑latency advantage for diffusion language models.

By Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum