Towards Data Science

Disaggregation Is a Thousand-GPU Problem

The article discusses the conditions under which separating the prefill phase from the decode phase in large language model inference is beneficial. It explains that only when three specific criteria are met does this split pay off, and it recommends using chunked prefill as the default approach when operating below a certain GPU threshold.

Hugging Face Trending Papers
Jun 9

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models

Diffusion large language models (dLLMs) re-encode the entire prefix at every denoising step, causing recomputation that scales quadratically with context length and becomes prohibitive for long-context scenarios. We propose Prefilling-dLLM, a training-free prefill-decode disaggregation framework for dLLMs that partitions the prefix into N chunks, caches their KV representations once, and selects the top-K most relevant chunks with intra-chunk token sparsity for decoding, showing that sparse prefilling can outperform dense attention while reducing per-step complexity from quadratic in the full sequence length to quadratic only in the decode length.

arXiv Computation and Language
Sep 1

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models

arXiv:2606.10537v2 Announce Type: replace Abstract: Diffusion large language models (dLLMs) re-encode the entire prefix at every denoising step, causing recomputation that scales quadratically with...

By Jing Xiong, Qi Han, Shansan Gong, Yunta Hsieh, Boyuan Zheng, Chengyue Wu, Chaofan Tao, Chenyang Zhao, Ngai Wong
arXiv Machine Learning
Sep 14

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

The paper investigates how a single SM utilization metric can misrepresent the true workload of large language model (LLM) inference on Nvidia Hopper GPUs. By profiling vLLM with FlashAttention‑3 and cuBLASLt on an H100 NVL across various phases (cold prefill, warm prefill, and decode) and varying sequence length and batch size, the authors replace the single utilization figure with eight detailed counter‑validated views. These views, tied to specific Nsight Compute counters or formulas, reveal how factors such as fragment fill, occupancy limits, stall signatures, wave quantization, and kernel selection create utilization gaps across four production models and six per‑layer kernel roles.

By Mohammad Siavashi, Gerald Q. Maguire Jr., Dejan Kostic, Marco Chiesa
arXiv AI
5d ago

Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits

The paper introduces Decode‑Latency Feedback Prefill (DLFP), a model‑free controller that adjusts prefill chunk sizes during concurrent autoregressive inference to reduce interference between new and ongoing requests. Implemented in vLLM, DLFP achieves significant reductions in P99 inter‑token latency on Qwen3‑0.6B while maintaining output correctness and SLO compliance, though it fails to generalize to larger models or multi‑GPU setups. The study highlights the limits of this approach and suggests the need for a completion‑timed controller for broader applicability.

By Gaurav Agarwal, Ashish Garg, Isha Singhal