arXiv:2412. 04504v2 Announce Type: replace-cross Abstract: As large language models (LLMs) grow in popularity for their diverse capabilities, improving the efficiency of their inference systems has become increasingly critical.
By Ozgur Guldogan, Jackson Kunde, Kangwook Lee, Ramtin Pedarsani
Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.
By Dohyun Park, Hubertus Franke, Daniel G. Waddington, Swaminathan Sundararaman, Yongjoo Park
arXiv:2607. 09172v1 Announce Type: cross Abstract: Large Language Models are reshaping how software is developed and maintained.
By Nada Zine, Tristan Coignion, Vincenzo Stoico, Cl\'ement Quinton, Romain Rouvoy, Patricia Lago
arXiv:2609.23130v1 Announce Type: new
Abstract: Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state,...
By Twinkll Sisodia
arXiv:2505. 07833v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) improves the reliability of large language models by integrating external knowledge, but serving RAG pipelines efficiently is challenging because requests traverse heterogeneous components spanning LLM inference, databases, and CPU-side processing.
By Saurabh Agarwal, Bodun Hu, Luis Pabon, Myungjin Lee, Jayanth Srinivasa, Aditya Akella
arXiv:2508. 06133v4 Announce Type: replace-cross Abstract: We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous prompt (prefill) and response (decode) lengths.
By Meixuan Wang, Yinyu Ye, Zijie Zhou
arXiv:2607. 18253v1 Announce Type: new Abstract: Modern language query routers improve inference efficiency by assigning each query to a model that balances response quality and monetary cost.
By Shivam Patel, Akaash R. Parthasarathy, Ankur Mallick, Gauri Joshi
arXiv:2607. 28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak.
By Jiaxuan Chen, Jianshu She, Ye Yuan, Rajat Ghosh, Karan Gupta, Qirong Ho, Xue Liu, Oana Balmau
arXiv:2510. 03243v3 Announce Type: replace-cross Abstract: Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that is becoming increasingly acute with the rise of reasoning-capable LLMs whose generation lengths are highly variable.
By Yiheng Tao, Yihe Zhang, Matthew Dearing, Xin Wang, Yuping Fan, Michael E. Papka, Zhiling Lan
arXiv:2609.05760v1 Announce Type: cross
Abstract: We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU enviro...
By Zlatan Feric, Amir Taherin, Bin Ren, Yanzhi Wang, Jennifer Dy, David Kaeli
AsyncFlow is an asynchronous streaming reinforcement learning framework designed to improve the post‑training phase of large language models. It introduces a distributed data storage and transfer module that enables panoramic data management and fine‑grained scheduling, allowing automated pipeline overlapping and dynamic load balancing. The framework also employs an asynchronous producer‑consumer workflow to reduce computational idleness by deferring parameter updates within staleness thresholds, and it is architecturally decoupled from training and inference engines, providing modular, customizable user interfaces. Experiments show an average throughput improvement of 1.59× over the state‑of‑the‑art baseline.
By Zhenyu Han, Ansheng You, Haibo Wang, Kui Luo, Guang Yang, Wenqi Shi, Menglong Chen, Sicheng Zhang, Zeshun Lan, Chunshi Deng, Huazhong Ji, Wenjie Liu, Yu Huang, Yixiang Zhang, Chenyi Pan, Jing Wang, Xin Huang, Chunsheng Li, Jianping Wu
arXiv:2608. 16336v1 Announce Type: cross Abstract: Modern LLM serving deployments must simultaneously satisfy heterogeneous service-level objectives (SLOs) across a diverse population of user tiers, ranging from latency-critical API calls to background batch processing.
By Anders Vestrum, Arya Raeesi, Hanna Roed