arXiv AI
Jun 12

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

arXiv:2505. 04021v3 Announce Type: replace-cross Abstract: Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increasingly important as token prices fall.

By Shan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li, Shuo Yang, Xinyuan Tong, Yang Wang, Zhiqiang Xie, Yuwei An, Shiyi Cao, Ke Bao, Deepak Vij, Xiaoning Ding, Yichen Wang, Qingda Lu, Zhong Wang, Gao Gao, Harry Xu, Junyi Shu, Jiarong Xing, Ying Sheng
arXiv Machine Learning
Sep 17

Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving

The paper introduces FairInference, a system that guarantees token-level latency isolation for well-behaved clients in multi-tenant LLM serving. It provides a δ-token fairness guarantee, ensuring that a token generated in isolation within time d will be produced within d + δ in a shared environment. The approach enforces per-token deadlines, bounds GPU compute sharing delays, and accounts for shared KV cache overhead, leading to reduced latency spikes and higher overall throughput compared to existing LLM serving systems.

By Dev Bali, Soujanya Ponnapalli, Yichuan Wang, Natacha Crooks, Scott Shenker, Matei Zaharia
arXiv AI
Jul 7

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.

By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral