arXiv Machine Learning By Ming-Yen Lee, Hanchen Yang, Faaiq Waqar, Harsono Simka, Tushar Krishna, Muhammed Ahosan Ul Karim, Shimeng Yu

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

Read the original on arXiv Machine Learning →

arXiv:2607. 26491v1 Announce Type: cross Abstract: The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management

The paper introduces LM‑CXD, a CXL‑SSD design tailored for large language model (LLM) prefix caching. By aligning KV chunk management between the serving engine and the storage device, exposing NAND-to‑DRAM progress, and using device DRAM as a GPU‑accessible buffer, LM‑CXD reduces time‑to‑first‑token (TTFT) by up to 4× compared to a stock CXL‑SSD and brings performance within 1.5× of local DRAM across five LLM models. The approach also incorporates windowed prefetching and layer‑wise KV movement to hide NAND latency under limited device DRAM.

By Hyunsun Chung, Taewan Noh, Minji Kim, Joo-Young Hwang, Hong-Yeon Kim, Youngjae Kim
arXiv Machine Learning
5d ago

The KV Cache Is the New Memory Wall

The paper argues that for large‑context autoregressive language‑model inference, memory bandwidth—specifically the Key‑Value (KV) cache—becomes the limiting resource rather than arithmetic throughput. It analytically derives how arithmetic intensity decays with context length for NVIDIA H100, NVIDIA B200, and AMD MI300X, identifies crossover points where KV traffic overtakes weight traffic, and evaluates representative techniques across five compression domains. The study finds a three‑regime behavior: below the crossover, weight traffic dominates and KV compression offers little benefit; beyond it, KV traffic dominates and compression methods trade quality for bandwidth, with paging and prefix sharing being lossless but capacity‑limited, while quantization and eviction directly reduce bandwidth at the cost of accuracy. whyItMatters":"The work provides a unified analytical framework and a standardized protocol that enable consistent comparison of KV‑compression techniques across hardware and workloads, guiding practitioners in selecting appropriate methods for long‑context inference."

By Tejinder Singh