arXiv AI By Dong Liu, Yanxuan Yu

MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators

Read the original on arXiv AI →

MeshKV is a Network‑on‑Chip key‑value cache fabric designed to improve transformer decoding on tiled accelerators. It uses packetized flows, TaKV affine striping, Mare multicast with duplicate suppression, and Pad to overlap prefetch, multiply, and softmax operations, thereby converting back‑pressure into useful KV transfer. In an 8×8 FPGA prototype with LLaMA‑2‑7B and Mistral‑7B, MeshKV cuts interconnect traffic by up to 58 %, doubles KV bandwidth utilization, and boosts multi‑stream throughput by up to 1.9×.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 29

The KV Cache Is the New Memory Wall

The paper argues that for large‑context autoregressive language‑model inference, memory bandwidth—specifically the Key‑Value (KV) cache—becomes the limiting resource rather than arithmetic throughput. It analytically derives how arithmetic intensity decays with context length for NVIDIA H100, NVIDIA B200, and AMD MI300X, identifies crossover points where KV traffic overtakes weight traffic, and evaluates representative techniques across five compression domains. The study finds a three‑regime behavior: below the crossover, weight traffic dominates and KV compression offers little benefit; beyond it, KV traffic dominates and compression methods trade quality for bandwidth, with paging and prefix sharing being lossless but capacity‑limited, while quantization and eviction directly reduce bandwidth at the cost of accuracy. whyItMatters":"The work provides a unified analytical framework and a standardized protocol that enable consistent comparison of KV‑compression techniques across hardware and workloads, guiding practitioners in selecting appropriate methods for long‑context inference."

By Tejinder Singh
arXiv Machine Learning
Jun 10

Operator Fusion for LLM Inference on the Tensix Architecture

arXiv:2606. 09879v1 Announce Type: new Abstract: This study addresses on-device inference bottlenecks of Transformer models on Tenstorrent's Tensix architecture and proposes an operator fusion strategy that enhances data locality.

By Qingbo Wu, Ke Li, Wenzhu Wang, Jie Yu, Ruian Zhang, Lili Liu