arXiv Machine Learning By Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet Pham, Shivaram Venkataraman

TeDiServe: High SLO Attainment Serving for Diffusion Language Models

Read the original on arXiv Machine Learning →

TeDiServe is a cluster‑level serving system designed for diffusion language models (DLMs). It addresses DLM‑specific challenges such as the speed‑quality tradeoff from confidence‑based denoising, variable parallelization under fluctuating load, and non‑uniform per‑step costs from approximate KV caching. By employing deadline‑aware scheduling, adaptive load control, and a quality‑aware optimization for cluster reconfiguration, TeDiServe achieves up to 56.6 percentage points higher SLO attainment and reduces end‑to‑end latency by up to 46% with less than 1% accuracy loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

The paper introduces the Denoising Workload Surface (DWS), a two‑dimensional probability surface that captures the block‑autoregressive generation structure of diffusion large language models (dLLMs). By preserving both output block and within‑block denoising step information, DWS enables a lightweight, prompt‑only predictor to estimate per‑request inference cost accurately, even on a single CPU core. In real‑world serving experiments, DWS reduces cost‑prediction error by up to 2.5× and improves end‑to‑end latency for online chatbots by up to 1.92×.

By Haoyu Zheng, Fangcheng Fu, Binhang Yuan, Yongqiang Zhang, Liang Deng, Hao Wang, Yuanyuan Zhu, Xiao Yan, Jiawei Jiang
arXiv Machine Learning
Aug 10

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.

By Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse
arXiv Machine Learning
Jul 13

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

arXiv:2607. 08930v1 Announce Type: new Abstract: Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency.

By Yuanjie Zhu, Liangwei Yang, Ke Xu, Weizhi Zhang, Shanghao Li, Zihe Song, Philip S. Yu