arXiv AI

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

arXiv:2511. 15927v4 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead.

arXiv Machine Learning
Sep 18

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any-order decoding and strong performance with parallel decoding.

By Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai
Hugging Face Trending Papers
Sep 17

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid AR backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any‑order decoding and strong performance under parallel decoding.

arXiv Machine Learning
Aug 6

Esoteric Language Models: A Family of Any-Order Diffusion LLMs

arXiv:2506. 01928v5 Announce Type: replace-cross Abstract: Diffusion-based language models offer a compelling alternative to autoregressive (AR) models by enabling parallel and controllable generation.

By Subham Sekhar Sahoo, Zhihan Yang, Yash Akhauri, Johnna Liu, Deepansha Singh, Zhoujun Cheng, Zhengzhong Liu, Eric Xing, John Thickstun, Arash Vahdat
arXiv Machine Learning
Jun 26

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM

arXiv:2606. 26120v1 Announce Type: cross Abstract: Diffusion Large Language Models (dLLMs) offer a promising alternative to autoregressive models, excelling in text generation tasks due to their bidirectional attention mechanisms.

By Tianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li, Yixin Chen, YunHao Cao, YiQi Hu, Zhuotao Tian
arXiv Machine Learning
Aug 4

A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs

arXiv:2603. 07475v4 Announce Type: replace-cross Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are trained through full-sequence denoising.

By Raghavv Goel, Risheek Garrepalli, Sudhanshu Agrawal, Chris Lott, Mingu Lee, Fatih Porikli
arXiv AI
Sep 3

Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

The paper introduces PILL, a new infilling technique for diffusion language models that eliminates the need for a preset initial length and reduces inference overhead. PILL uses probing-based length-free decoding, cutting down on extra forward passes and speeding up generation. Experiments across five diffusion models and eight benchmarks show PILL outperforms the strongest baseline with higher pass rates and BLEU-2 scores while running 1.82× faster.

By Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong