arXiv Machine Learning By Aman Sunesh, Ali Alshehhi, Hivansh Dhakne

RequestRouter: Request-Boundary Routing for Efficient Single-GPU LLM Inference

Read the original on arXiv Machine Learning →

arXiv:2605. 23057v2 Announce Type: replace Abstract: RequestRouter is a lightweight request-boundary controller for reducing the latency and energy cost of single-GPU large language model inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 16

Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving

The paper evaluates a learned request‑routing policy for disaggregated large‑language‑model serving, where compute‑heavy prefill and memory‑heavy decode stages run on separate GPU pools. Using a discrete‑event simulator and real NVIDIA A40 GPUs, the calibrated router—leveraging prompt length, predicted output length, KV‑cache pressure, and SLO class—outperforms round‑robin, least‑loaded, and length‑based heuristics, achieving the highest mean goodput (0.864) and lowest variance across three mixed, bursty arrival traces. Hardware calibration proves critical, providing a 4.5‑point goodput boost and roughly 40 % of the tail‑latency advantage, and the learned router can match round‑robin performance with one fewer GPU in certain scenarios.

By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv Machine Learning
Sep 18

PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

PrefixBench-H100 is a reproducible benchmark that evaluates how reusing prompt prefixes affects LLM serving performance on NVIDIA H100 GPUs. It tests two popular runtimes (vLLM and TensorRT-LLM) across varied workloads, measuring metrics such as time-to-first-token, latency, throughput, cache hits, and GPU memory usage. The study identifies when prefix reuse significantly reduces first‑token latency and when cache pressure diminishes those gains, noting that cache effectiveness is largely unaffected by concurrency or output length, while differences arise mainly in scheduling.

By Omkar Shewale, Deepak Kumar, Divakar Kumar Yadav
arXiv AI
Jun 30

KernelSight-LM: A Kernel-Level LLM Inference Simulator

arXiv:2606. 28565v1 Announce Type: cross Abstract: As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency targets.

By Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt