arXiv Machine Learning By Nitin Kedia, Saurabh Agarwal, Myungjin Lee, Aditya Akella

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

Read the original on arXiv Machine Learning →

arXiv:2607. 04206v1 Announce Type: cross Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 10

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.

By Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse
arXiv Machine Learning
1d ago

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

The paper introduces the Denoising Workload Surface (DWS), a two‑dimensional probability surface that captures the block‑autoregressive generation structure of diffusion large language models (dLLMs). By preserving both output block and within‑block denoising step information, DWS enables a lightweight, prompt‑only predictor to estimate per‑request inference cost accurately, even on a single CPU core. In real‑world serving experiments, DWS reduces cost‑prediction error by up to 2.5× and improves end‑to‑end latency for online chatbots by up to 1.92×.

By Haoyu Zheng, Fangcheng Fu, Binhang Yuan, Yongqiang Zhang, Liang Deng, Hao Wang, Yuanyuan Zhu, Xiao Yan, Jiawei Jiang
arXiv Machine Learning
Jul 13

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

arXiv:2607. 08930v1 Announce Type: new Abstract: Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency.

By Yuanjie Zhu, Liangwei Yang, Ke Xu, Weizhi Zhang, Shanghao Li, Zihe Song, Philip S. Yu