arXiv Machine Learning By Dev Bali, Soujanya Ponnapalli, Yichuan Wang, Natacha Crooks, Scott Shenker, Matei Zaharia

Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving

Read the original on arXiv Machine Learning →

The paper introduces FairInference, a system that guarantees token-level latency isolation for well-behaved clients in multi-tenant LLM serving. It provides a δ-token fairness guarantee, ensuring that a token generated in isolation within time d will be produced within d + δ in a shared environment. The approach enforces per-token deadlines, bounds GPU compute sharing delays, and accounts for shared KV cache overhead, leading to reduced latency spikes and higher overall throughput compared to existing LLM serving systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 10

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.

By Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse