How Long Prompts Block Other Requests - Optimizing LLM Performance
Related stories
Prefill and Decode for Concurrent Requests - Optimizing LLM Performance
Optimizing your LLM in production
Multi-Bin Batching for Increasing LLM Inference Throughput
arXiv:2412. 04504v2 Announce Type: replace-cross Abstract: As large language models (LLMs) grow in popularity for their diverse capabilities, improving the efficiency of their inference systems has become increasingly critical.
Jumping the Line: Exploiting Length Predictions in LLM Scheduling
arXiv:2610.03430v1 Announce Type: new Abstract: Efficient request scheduling is increasingly important for reducing completion time in large language model (LLM) serving. Size-based policies such as...
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving
arXiv:2608. 08382v1 Announce Type: new Abstract: As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control.
Scorpio: Serving Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
Scorpio is an LLM serving system that optimizes for heterogeneous Service Level Objectives (SLOs) such as Time to First Token (TTFT) and Time Per Output Token (TPOT). It uses adaptive scheduling across admission control, queue management, and batch selection, featuring a TTFT Guard that reorders requests by least-deadline-first and rejects unattainable ones, and a TPOT Guard that employs VBS-based admission control and a credit-based batching mechanism. Predictive modules support both guards, and evaluations show Scorpio can increase system goodput by up to 14.4× and improve SLO adherence by up to 46.5% under high load compared to state-of-the-art baselines.
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving
arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.
LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs
arXiv:2608. 08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs).
LLM Inference Under Bursty Workload Distribution: Modifying the WAIT Algorithm
arXiv:2608. 06135v1 Announce Type: new Abstract: Large Language Models (LLMs) such as ChatGPT and Claude are widely used for information retrieval and problem-solving.
Evaluating Test-Time Scaling of General LLM Agents
arXiv:2602.18998v2 Announce Type: replace Abstract: LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavio...