Beyond Prediction: Tail-Aware Scheduling for LLM Inference
arXiv:2606. 18431v1 Announce Type: new Abstract: LLM serving exhibits extreme length variability, making size-based scheduling difficult in practice.
arXiv:2608. 06135v1 Announce Type: new Abstract: Large Language Models (LLMs) such as ChatGPT and Claude are widely used for information retrieval and problem-solving.
arXiv:2606. 18431v1 Announce Type: new Abstract: LLM serving exhibits extreme length variability, making size-based scheduling difficult in practice.
arXiv:2607. 18253v1 Announce Type: new Abstract: Modern language query routers improve inference efficiency by assigning each query to a model that balances response quality and monetary cost.
arXiv:2609.17193v1 Announce Type: new Abstract: Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed...
arXiv:2412. 04504v2 Announce Type: replace-cross Abstract: As large language models (LLMs) grow in popularity for their diverse capabilities, improving the efficiency of their inference systems has become increasingly critical.
arXiv:2510. 03243v3 Announce Type: replace-cross Abstract: Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that is becoming increasingly acute with the rise of reasoning-capable LLMs whose generation lengths are highly variable.
arXiv:2508. 06133v4 Announce Type: replace-cross Abstract: We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous prompt (prefill) and response (decode) lengths.
The paper investigates how different coordination topologies—sequential, star, and full-mesh—affect traffic patterns in multi-agent large language model (LLM) systems. By measuring LLM-call interarrival times over 500 runs per topology, the authors find that topology shapes request arrival processes, with fan-out coordination producing a bimodal distribution and the reasoning phase following a log-normal distribution rather than a Poisson model. These structural differences influence inference and network-level metrics.
Scorpio is an LLM serving system that optimizes for heterogeneous Service Level Objectives (SLOs) such as Time to First Token (TTFT) and Time Per Output Token (TPOT). It uses adaptive scheduling across admission control, queue management, and batch selection, featuring a TTFT Guard that reorders requests by least-deadline-first and rejects unattainable ones, and a TPOT Guard that employs VBS-based admission control and a credit-based batching mechanism. Predictive modules support both guards, and evaluations show Scorpio can increase system goodput by up to 14.4× and improve SLO adherence by up to 46.5% under high load compared to state-of-the-art baselines.
arXiv:2607. 19349v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge.
arXiv:2606. 00946v1 Announce Type: cross Abstract: Efficiently serving large language model (LLM) inference tasks is crucial both for user-perceived latency such as time-to-first-token (TTFT) and for GPU utilization.
TeDiServe is a cluster‑level serving system designed for diffusion language models (DLMs). It addresses DLM‑specific challenges such as the speed‑quality tradeoff from confidence‑based denoising, variable parallelization under fluctuating load, and non‑uniform per‑step costs from approximate KV caching. By employing deadline‑aware scheduling, adaptive load control, and a quality‑aware optimization for cluster reconfiguration, TeDiServe achieves up to 56.6 percentage points higher SLO attainment and reduces end‑to‑end latency by up to 46% with less than 1% accuracy loss.