arXiv Machine Learning

FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration

arXiv:2608. 19659v1 Announce Type: new Abstract: Choosing tensor-parallel (TP) degrees and replica counts for an LLM serving fleet is difficult because performance is not monotonic in TP and the feasible choice can change with load.

arXiv AI
Sep 16

Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving

The paper evaluates a learned request‑routing policy for disaggregated large‑language‑model serving, where compute‑heavy prefill and memory‑heavy decode stages run on separate GPU pools. Using a discrete‑event simulator and real NVIDIA A40 GPUs, the calibrated router—leveraging prompt length, predicted output length, KV‑cache pressure, and SLO class—outperforms round‑robin, least‑loaded, and length‑based heuristics, achieving the highest mean goodput (0.864) and lowest variance across three mixed, bursty arrival traces. Hardware calibration proves critical, providing a 4.5‑point goodput boost and roughly 40 % of the tail‑latency advantage, and the learned router can match round‑robin performance with one fewer GPU in certain scenarios.

By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv AI
Sep 24

Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Crossflow introduces an elastic boundary for prefilling and decoding in large language model serving, allowing decode nodes to publish short‑lived leases that limit prefilling resources and output projections. By adapting to dynamic phase demand, Crossflow improves token throughput by 16.2‑17.4% on average and up to 43.4% under high load, while consistently reducing mean time‑to‑first‑token. The approach eliminates the inefficiencies of static partitioning, which can leave 17% of cluster capacity idle or cause queueing and lost throughput.

By Yi Xu, Ehsan K. Ardestani, Wenyin Fu, Martin Schatz, Krishna Malladi, Zhan Shu, Adnan Aziz, Shobhit Kanaujia, Ajit Mathews, Chunqiang Tang
arXiv AI
Jun 30

KernelSight-LM: A Kernel-Level LLM Inference Simulator

arXiv:2606. 28565v1 Announce Type: cross Abstract: As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency targets.

By Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt