TeDiServe is a cluster‑level serving system designed for diffusion language models (DLMs). It addresses DLM‑specific challenges such as the speed‑quality tradeoff from confidence‑based denoising, variable parallelization under fluctuating load, and non‑uniform per‑step costs from approximate KV caching. By employing deadline‑aware scheduling, adaptive load control, and a quality‑aware optimization for cluster reconfiguration, TeDiServe achieves up to 56.6 percentage points higher SLO attainment and reduces end‑to‑end latency by up to 46% with less than 1% accuracy loss.
By Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet Pham, Shivaram Venkataraman
arXiv:2607. 12829v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models.
By Daehoon Gwak, Minhyung Lee, Junwoo Park, Jaegul Choo
The paper introduces the Denoising Workload Surface (DWS), a two‑dimensional probability surface that captures the block‑autoregressive generation structure of diffusion large language models (dLLMs). By preserving both output block and within‑block denoising step information, DWS enables a lightweight, prompt‑only predictor to estimate per‑request inference cost accurately, even on a single CPU core. In real‑world serving experiments, DWS reduces cost‑prediction error by up to 2.5× and improves end‑to‑end latency for online chatbots by up to 1.92×.
By Haoyu Zheng, Fangcheng Fu, Binhang Yuan, Yongqiang Zhang, Liang Deng, Hao Wang, Yuanyuan Zhu, Xiao Yan, Jiawei Jiang
arXiv:2607. 08930v1 Announce Type: new Abstract: Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency.
By Yuanjie Zhu, Liangwei Yang, Ke Xu, Weizhi Zhang, Shanghao Li, Zihe Song, Philip S. Yu
Flash-dLLM is a training‑free inference acceleration framework that improves the speed and memory efficiency of Diffusion Large Language Models (dLLMs). It tackles GPU memory I/O bottlenecks by introducing an I/O‑aware fused KV‑cache kernel and then employs a draft‑and‑verify decoding strategy that uses the dLLM itself as both drafter and verifier. Experiments on mathematical reasoning and code‑generation tasks show Flash‑dLLM outperforms existing acceleration methods, achieving up to 11.0× speedups over the Elastic‑Cache baseline.
By Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
arXiv:2607. 04206v1 Announce Type: cross Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation.
By Nitin Kedia, Saurabh Agarwal, Myungjin Lee, Aditya Akella