arXiv AI

Evaluation of Baseline Methods for IDD-based SSD External Memory Search

arXiv:2606. 01840v1 Announce Type: new Abstract: Many difficult search problems cannot be solved by algorithms such as A* using only RAM.

arXiv AI
Sep 24

Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management

The paper introduces LM‑CXD, a CXL‑SSD design tailored for large language model (LLM) prefix caching. By aligning KV chunk management between the serving engine and the storage device, exposing NAND-to‑DRAM progress, and using device DRAM as a GPU‑accessible buffer, LM‑CXD reduces time‑to‑first‑token (TTFT) by up to 4× compared to a stock CXL‑SSD and brings performance within 1.5× of local DRAM across five LLM models. The approach also incorporates windowed prefetching and layer‑wise KV movement to hide NAND latency under limited device DRAM.

By Hyunsun Chung, Taewan Noh, Minji Kim, Joo-Young Hwang, Hong-Yeon Kim, Youngjae Kim
arXiv Machine Learning
Sep 25

When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse

The paper investigates cache replacement strategies for large language model (LLM) prefix reuse, analyzing production traces from two companies and testing 14 eviction algorithms in both high-bandwidth memory (HBM) and large memory-pool environments. It finds that sophisticated policies designed for traditional caches offer little advantage over simple LRU, because prefix reuse is largely driven by the regular pacing of active sessions, making recency a strong predictor. The study also highlights new challenges such as heavy-tailed session footprints and variable miss costs, and proposes a compute-savings ratio along with two offline oracles to better quantify these effects, suggesting that effective prefix-cache management should combine recency with selective quick demotion, compute-aware partial eviction, and capacity-dependent granularity.

By Yiyu Liu, Minlan Yu, Juncheng Yang
arXiv Machine Learning
5d ago

Evaluating the accuracy of KV cache reuse techniques

The paper discusses position‑independent KV cache reuse, a technique designed to cut latency in retrieval‑augmented generation by reusing chunk‑level KV caches across prompts. It argues that current evaluation methods overstate the accuracy of such reuse because they do not accurately capture the loss of accuracy, and that existing datasets lack the necessary reuse dynamics for thorough testing. To remedy this, the authors propose a new evaluation methodology that unambiguously measures accuracy loss and introduce Boxoffice, a tool that programmatically creates datasets with challenging KV cache reuse patterns.

By Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona
arXiv AI
Jun 19

Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices

arXiv:2606. 19528v1 Announce Type: cross Abstract: Fine-tuning of Large Language Models (LLMs) using Low-Rank Adaptation (LoRA) on an end-user's data offers personalized experiences while keeping data private, but faces severe memory constraints on consumer hardware.

By Hassan Dbouk, Matthias Reisser, Prathamesh Mandke, Likhita Arun Navali, Christos Louizos