arXiv:2609.13585v1 Announce Type: cross
Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...
By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv:2607. 28633v1 Announce Type: cross Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly.
By Sanjeev Rao Ganjihal
Scepsy is a serving system designed to efficiently schedule arbitrary multi‑LLM agentic workflows on GPU clusters. It leverages the observation that each LLM’s share of execution time remains relatively stable across requests, profiling LLMs under various parallelism levels to build an Aggregate LLM Pipeline that predicts throughput and latency. Using this predictor, Scepsy searches for optimal GPU allocations—balancing fractional GPU shares, tensor parallelism, and replica counts—and then heuristically places them on the cluster to reduce fragmentation and honor network topology, achieving up to 2.5× higher throughput and 1.0–3.3× lower latency compared to baseline approaches.
By Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhao, Yanda Tao, Pedro Silvestre, Guo Li, Huanzhou Zhu, Llu\'is Vilanova, Peter Pietzuch
arXiv:2504. 08791v3 Announce Type: replace-cross Abstract: On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability.
By Zonghang Li, Tao Li, Wenjiao Feng, Rongxing Xiao, Jianshu She, Hong Huang, Mohsen Guizani, Hongfang Yu, Qirong Ho, Wei Xiang, Xue Liu
arXiv:2512. 10236v2 Announce Type: replace-cross Abstract: Modern ML workloads demand distributing training and inference across multiple GPUs.
By Shagnik Pal, Shaizeen Aga, Suchita Pati, Mahzabeen Islam, Lizy K. John
arXiv:2607. 23264v1 Announce Type: cross Abstract: Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation.
By Jianwen Xian, Zhiyuan Xu, Yuchen Li, Ziliang Lai, Kang He, Zhen Huang, Aichen Feng, Jinyan Chen, Yilin Zhang, Qinqin Chen, Chengru Song
arXiv:2511. 04791v2 Announce Type: replace Abstract: Modern LLM serving systems must sustain high throughput while meeting strict latency SLOs across two distinct inference phases: compute-intensive prefill and memory-bound decode phases.
By Lei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong, Mark Hill, Murali Annavaram
arXiv:2607. 19539v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes.
By Minyu Cui, Anna Wingkvist, Morgan Ericsson
arXiv:2609.34818v2 Announce Type: replace-cross
Abstract: AlphaFold2 is written in JAX, so the same inference code compiles and runs unchanged on CPUs, GPUs and Google Cloud TPUs. That portability ma...
By Lorenzo Pazienza, Ihab El Bani
The paper argues that for large‑context autoregressive language‑model inference, memory bandwidth—specifically the Key‑Value (KV) cache—becomes the limiting resource rather than arithmetic throughput. It analytically derives how arithmetic intensity decays with context length for NVIDIA H100, NVIDIA B200, and AMD MI300X, identifies crossover points where KV traffic overtakes weight traffic, and evaluates representative techniques across five compression domains. The study finds a three‑regime behavior: below the crossover, weight traffic dominates and KV compression offers little benefit; beyond it, KV traffic dominates and compression methods trade quality for bandwidth, with paging and prefix sharing being lossless but capacity‑limited, while quantization and eviction directly reduce bandwidth at the cost of accuracy.
whyItMatters":"The work provides a unified analytical framework and a standardized protocol that enable consistent comparison of KV‑compression techniques across hardware and workloads, guiding practitioners in selecting appropriate methods for long‑context inference."
By Tejinder Singh
arXiv:2606. 28565v1 Announce Type: cross Abstract: As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency targets.
By Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt
arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.
By Xinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji, Jinchen Jiang, Zepeng Zhao, Ziruo Xiao, Zihao Ye, Yingyi Huang, Ruihang Lai, Hongyi Jin, Bohan Hou, Mengdi Wu, Yixin Dong, Anthony Yip, Zihao Ye, Songting Wang, Wenqin Yang, Xupeng Miao, Tianqi Chen, Zhihao Jia