arXiv AI By Sanjeev Rao Ganjihal

Topology-Aware Data Movement for Disaggregated GPU Inference

Read the original on arXiv AI →

arXiv:2607. 28633v1 Announce Type: cross Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

arXiv:2609.13585v1 Announce Type: cross Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...

By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv Machine Learning
5d ago

The KV Cache Is the New Memory Wall

The paper argues that for large‑context autoregressive language‑model inference, memory bandwidth—specifically the Key‑Value (KV) cache—becomes the limiting resource rather than arithmetic throughput. It analytically derives how arithmetic intensity decays with context length for NVIDIA H100, NVIDIA B200, and AMD MI300X, identifies crossover points where KV traffic overtakes weight traffic, and evaluates representative techniques across five compression domains. The study finds a three‑regime behavior: below the crossover, weight traffic dominates and KV compression offers little benefit; beyond it, KV traffic dominates and compression methods trade quality for bandwidth, with paging and prefix sharing being lossless but capacity‑limited, while quantization and eviction directly reduce bandwidth at the cost of accuracy. whyItMatters":"The work provides a unified analytical framework and a standardized protocol that enable consistent comparison of KV‑compression techniques across hardware and workloads, guiding practitioners in selecting appropriate methods for long‑context inference."

By Tejinder Singh