arXiv Machine Learning By Yuanjie Zhu, Liangwei Yang, Ke Xu, Weizhi Zhang, Shanghao Li, Zihe Song, Philip S. Yu

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

Read the original on arXiv Machine Learning →

arXiv:2607. 08930v1 Announce Type: new Abstract: Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 9

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency. We present BlockServe, a continuous batching framework that integrates block-grained scheduling -- immediately evicting completed requests at block boundaries -- with mixed-state execution that extends dual cache and parallel decoding to heterogeneous batches via gather-scatter indexing.

arXiv Machine Learning
5d ago

TeDiServe: High SLO Attainment Serving for Diffusion Language Models

TeDiServe is a cluster‑level serving system designed for diffusion language models (DLMs). It addresses DLM‑specific challenges such as the speed‑quality tradeoff from confidence‑based denoising, variable parallelization under fluctuating load, and non‑uniform per‑step costs from approximate KV caching. By employing deadline‑aware scheduling, adaptive load control, and a quality‑aware optimization for cluster reconfiguration, TeDiServe achieves up to 56.6 percentage points higher SLO attainment and reduces end‑to‑end latency by up to 46% with less than 1% accuracy loss.

By Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet Pham, Shivaram Venkataraman
arXiv AI
Aug 26

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

Masked diffusion language models (dLLMs) promise faster text generation by denoising multiple tokens simultaneously, yet their real‑world serving behavior has been largely unexamined. Using LLaDA‑8B‑Instruct on a single NVIDIA H200 GPU, the study finds that request difficulty is discretized into 11 step‑count levels, short‑budget benchmarks underestimate serving variance, and only 24% of single‑request time is GPU computation, with batching mainly reducing CPU dispatch overhead. The authors also demonstrate that output quality remains stable across batch sizes and propose a batch‑timeout rule for synchronized batching under Poisson arrivals.

By Farhana Amin, Sabiha Afroz, Mona Moghadampanah, Dimitrios S. Nikolopoulos