arXiv:2607. 02805v1 Announce Type: cross Abstract: High-throughput long-context generation is one of the central challenges for large language models.
By Pranshu Chaturvedi, Parth Shroff, Tarun Suresh, Hangoo Kang, Kaiyue Wen
arXiv:2606. 19475v1 Announce Type: new Abstract: Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range of tasks.
By Thomas Bertolani, Davide Bucciarelli, Leonardo Zini, Marcella Cornia, Lorenzo Baraldi
The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any-order decoding and strong performance with parallel decoding.
By Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai
arXiv:2606. 01774v1 Announce Type: cross Abstract: Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck for low-latency deployment.
By Yuchen Zhu, Jing Shi, Chongjian Ge, Hao Tan, Yiran Xu, Wanrong Zhu, Jason Kuen, Koustava Goswami, Rajiv Jain, Yongxin Chen, Molei Tao, Jiuxiang Gu
The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid AR backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any‑order decoding and strong performance under parallel decoding.
arXiv:2508. 10875v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm.
By Tianyi Li, Mingda Chen, Bowei Guo, Zhiqiang Shen