arXiv Machine Learning By Yijie Jin, Jiajun Xu, Yuxuan Liu, Chenkai Xu, Yi Tu, Jiajun Li, Dandan Tu, Xiaohui Yan, Kai Yu, Pengfei Liu, Zhijie Deng

Multi-Block Diffusion Language Models

Read the original on arXiv Machine Learning →

arXiv:2606. 29215v1 Announce Type: new Abstract: Block Diffusion Language Models (BD-LMs) improve diffusion-based text generation with KV caching and flexible-length generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

The paper introduces Block Parallelism (BP) and Context‑Sharded Block Parallelism (CSBP) to improve training efficiency for Block Diffusion Language Models (BDLMs) with long contexts. By assigning each corrupted‑block computation to a separate rank and sharding the shared clean sequence, CSBP reduces communication overhead and memory usage while preserving training semantics. Experiments on 16 H200 GPUs and 8 H100 GPUs show throughput gains of up to 1.61× and 7.59×, respectively, and higher benchmark pass rates in practical fine‑tuning scenarios.

By Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini
Hugging Face Trending Papers
Jul 20

FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models

Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly serial. Prior work has attempted to unlock inter-block parallelism through post-training methods, but achieves only modest speedups and often degrades accuracy.

arXiv AI
Jul 28

Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

arXiv:2607. 22663v1 Announce Type: new Abstract: Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic attention cost tractable.

By Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du