Hugging Face Trending Papers

Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing approaches use a single network for both context representation and iterative denoising, forcing one model to serve both roles and limiting its capacity for either role.

arXiv AI
Jul 20

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

arXiv:2511. 15927v4 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead.

By Vaibhav Singh, Oleksiy Ostapenko, Pierre-Andr\'e No\"el, Eugene Belilovsky, Torsten Scholak
arXiv AI
Jun 15

Residual Context Diffusion Language Models

arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.

By Yuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi, Coleman Hooper, Jintao Zhang, Aditya Tomar, Michael W. Mahoney, Sewon Min, Mehrdad Farajtabar, Kurt Keutzer, Amir Gholami, Chenfeng Xu
arXiv Machine Learning
Sep 14

Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

The paper investigates how block‑diffusion language models can use a constant‑size cache to enable efficient parallel decoding. By employing sequence mixers that summarize completed blocks into a reusable state and a block‑causal training objective, the authors pretrain three 3B block‑diffusion denoisers (attention, Mamba, and hybrid) on 300 B tokens. The resulting state‑space cache remains O(1) in memory and latency regardless of context length, yielding significant speed‑up and memory savings compared to traditional attention‑based caches, especially at very long sequences.

By Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko
Hugging Face Trending Papers
Aug 2

DeltaFlow: Noise-Adaptive Bidirectional Gated Delta Networks for Embedded Language Flows

Embedded Language Flows (ELF) rely primarily on full non-causal attention for iterative denoising, repeatedly incurring quadratic sequence-mixing cost at each sampling step. Gated Delta Networks (GDNs) provide an efficient recurrent alternative, but their standard causal formulation cannot directly capture the bidirectional context required by ELF.

arXiv Machine Learning
Sep 18

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any-order decoding and strong performance with parallel decoding.

By Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai
Hugging Face Trending Papers
Sep 17

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid AR backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any‑order decoding and strong performance under parallel decoding.

arXiv Machine Learning
Sep 18

Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

The paper introduces Block Parallelism (BP) and Context‑Sharded Block Parallelism (CSBP) to improve training efficiency for Block Diffusion Language Models (BDLMs) with long contexts. By assigning each corrupted‑block computation to a separate rank and sharding the shared clean sequence, CSBP reduces communication overhead and memory usage while preserving training semantics. Experiments on 16 H200 GPUs and 8 H100 GPUs show throughput gains of up to 1.61× and 7.59×, respectively, and higher benchmark pass rates in practical fine‑tuning scenarios.

By Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini