arXiv AI By Vaibhav Singh, Oleksiy Ostapenko, Pierre-Andr\'e No\"el, Eugene Belilovsky, Torsten Scholak

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

Read the original on arXiv AI →

arXiv:2511. 15927v4 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any-order decoding and strong performance with parallel decoding.

By Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai
Hugging Face Trending Papers
Sep 17

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid AR backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any‑order decoding and strong performance under parallel decoding.