arXiv AI

Online Linear Programming for Multi-Objective Routing in LLM Serving

arXiv:2607. 03948v1 Announce Type: new Abstract: We study the online routing problem in large language model serving, where requests arrive sequentially and must be dispatched to parallel decode workers under tight batch-size and KV-cache constraints.

arXiv AI
4d ago

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

arXiv:2605.18859v3 Announce Type: replace-cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single u...

By Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing, Wentao Guo, Yuhang Yao, Yuhang Han, Hanchen Li, Xu Wang, Zeyu Wang, Jie Xiao, Anjie Yang, Liang Tian, Lynn Ai, Eric Yang, Tianyu Shi
arXiv Machine Learning
Aug 10

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.

By Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse
arXiv Machine Learning
5d ago

TeDiServe: High SLO Attainment Serving for Diffusion Language Models

TeDiServe is a cluster‑level serving system designed for diffusion language models (DLMs). It addresses DLM‑specific challenges such as the speed‑quality tradeoff from confidence‑based denoising, variable parallelization under fluctuating load, and non‑uniform per‑step costs from approximate KV caching. By employing deadline‑aware scheduling, adaptive load control, and a quality‑aware optimization for cluster reconfiguration, TeDiServe achieves up to 56.6 percentage points higher SLO attainment and reduces end‑to‑end latency by up to 46% with less than 1% accuracy loss.

By Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet Pham, Shivaram Venkataraman