Prefill and Decode for Concurrent Requests - Optimizing LLM Performance
Related stories
How Long Prompts Block Other Requests - Optimizing LLM Performance
Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving
arXiv:2607. 02043v1 Announce Type: cross Abstract: Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering.
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving
arXiv:2608. 08382v1 Announce Type: new Abstract: As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control.
Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation. Their bidirectional attention prevents exact autoregressive-style KV caching, since committing one position shifts the KV activations of all others.
Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving
Crossflow introduces an elastic boundary for prefilling and decoding in large language model serving, allowing decode nodes to publish short‑lived leases that limit prefilling resources and output projections. By adapting to dynamic phase demand, Crossflow improves token throughput by 16.2‑17.4% on average and up to 43.4% under high load, while consistently reducing mean time‑to‑first‑token. The approach eliminates the inefficiencies of static partitioning, which can leave 17% of cluster capacity idle or cause queueing and lost throughput.
Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
arXiv:2607. 04206v1 Announce Type: cross Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation.
Threshold-Based Exclusive Batching for LLM Inference
arXiv:2606. 00516v1 Announce Type: new Abstract: Mixed batching (MB)--interleaving prefill and decode in a single batch--has become the standard scheduling strategy for large language model (LLM) inference due to its efficiency in maximizing compute and memory utilization.
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving
arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.
Multi-Bin Batching for Increasing LLM Inference Throughput
arXiv:2412. 04504v2 Announce Type: replace-cross Abstract: As large language models (LLMs) grow in popularity for their diverse capabilities, improving the efficiency of their inference systems has become increasingly critical.
Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
The paper introduces Decode‑Latency Feedback Prefill (DLFP), a model‑free controller that adjusts prefill chunk sizes during concurrent autoregressive inference to reduce interference between new and ongoing requests. Implemented in vLLM, DLFP achieves significant reductions in P99 inter‑token latency on Qwen3‑0.6B while maintaining output correctness and SLO compliance, though it fails to generalize to larger models or multi‑GPU setups. The study highlights the limits of this approach and suggests the need for a completion‑timed controller for broader applicability.
Overcoming the Communication-Performance Tradeoff in LLM Pretraining
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.