Hugging Face Blog

How Long Prompts Block Other Requests - Optimizing LLM Performance

arXiv Machine Learning
Aug 27

Scorpio: Serving Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference

Scorpio is an LLM serving system that optimizes for heterogeneous Service Level Objectives (SLOs) such as Time to First Token (TTFT) and Time Per Output Token (TPOT). It uses adaptive scheduling across admission control, queue management, and batch selection, featuring a TTFT Guard that reorders requests by least-deadline-first and rejects unattainable ones, and a TPOT Guard that employs VBS-based admission control and a credit-based batching mechanism. Predictive modules support both guards, and evaluations show Scorpio can increase system goodput by up to 14.4× and improve SLO adherence by up to 46.5% under high load compared to state-of-the-art baselines.

By Yinghao Tang, Tingfeng Lan, Bo Pan, Xiuqi Huang, Hui Lu, Wei Chen
arXiv Machine Learning
Aug 10

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.

By Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse