Architecting cost-effective infrastructure by navigating the latency and storage trade-offs of HNSW, SPANN, and DiskANN The post How to Optimize Vector Search When RAM Gets Too Expensive: On-Disk vs. In-Memory ANN Indexes appeared first on Towards Data Science .
By Oleg Tereshin
The paper introduces LM‑CXD, a CXL‑SSD design tailored for large language model (LLM) prefix caching. By aligning KV chunk management between the serving engine and the storage device, exposing NAND-to‑DRAM progress, and using device DRAM as a GPU‑accessible buffer, LM‑CXD reduces time‑to‑first‑token (TTFT) by up to 4× compared to a stock CXL‑SSD and brings performance within 1.5× of local DRAM across five LLM models. The approach also incorporates windowed prefetching and layer‑wise KV movement to hide NAND latency under limited device DRAM.
By Hyunsun Chung, Taewan Noh, Minji Kim, Joo-Young Hwang, Hong-Yeon Kim, Youngjae Kim
The paper investigates cache replacement strategies for large language model (LLM) prefix reuse, analyzing production traces from two companies and testing 14 eviction algorithms in both high-bandwidth memory (HBM) and large memory-pool environments. It finds that sophisticated policies designed for traditional caches offer little advantage over simple LRU, because prefix reuse is largely driven by the regular pacing of active sessions, making recency a strong predictor. The study also highlights new challenges such as heavy-tailed session footprints and variable miss costs, and proposes a compute-savings ratio along with two offline oracles to better quantify these effects, suggesting that effective prefix-cache management should combine recency with selective quick demotion, compute-aware partial eviction, and capacity-dependent granularity.
By Yiyu Liu, Minlan Yu, Juncheng Yang
arXiv:2605. 26168v2 Announce Type: replace-cross Abstract: Any device that runs Linux uses the Linux page cache, a central pillar in OS and application performance, serving to reduce extraneous disk access.
By Zejia Qi
arXiv:2608.22141v1 Announce Type: new
Abstract: Enterprise repositories are large, heteroge- neous, and continuously updated, making re- trieval difficult when efficient access, source- faithful evid...
By Xinyuan Song, Bowen Zhu, Hasibul Haque, Liang Zhao
arXiv:2606. 31121v1 Announce Type: new Abstract: Sequentially evolving LLM memory enables agents to reuse past experience, but existing systems usually deploy each locally generated memory update without checking whether it improves future behavior.
By Zihan Chen, Songwei Dong, Chengshuai Shi, Peng Wang, Song Wang, Cong Shen, Jundong Li
arXiv:2609.20276v2 Announce Type: replace-cross
Abstract: Memoizing an expensive function of a sorted score vector is a data-structure problem before it is a numerical one: at a billion gridpoints, a...
By Tamal Maharaj
arXiv:2607. 23632v1 Announce Type: cross Abstract: Science-intensive data profiling focuses on discovery and validation of various patterns in datasets.
By Yakov Kuzin, Dmitriy Shcheka, Michael Polyntsov, Kirill Stupakov, Mikhail Firsov, George Chernishev
The paper discusses position‑independent KV cache reuse, a technique designed to cut latency in retrieval‑augmented generation by reusing chunk‑level KV caches across prompts. It argues that current evaluation methods overstate the accuracy of such reuse because they do not accurately capture the loss of accuracy, and that existing datasets lack the necessary reuse dynamics for thorough testing. To remedy this, the authors propose a new evaluation methodology that unambiguously measures accuracy loss and introduce Boxoffice, a tool that programmatically creates datasets with challenging KV cache reuse patterns.
By Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona
arXiv:2608.30647v1 Announce Type: cross
Abstract: Language models can answer from precomputed memory, a model's saved reading of a body of material, reused across requests instead of read again at ea...
By Asa Shepard
arXiv:2209. 00188v4 Announce Type: replace-cross Abstract: Long-latency load requests continue to limit the performance of high-performance processors.
By Rahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo, Ataberk Olgun, Mohammad Sadrosadati, Onur Mutlu
arXiv:2606. 19528v1 Announce Type: cross Abstract: Fine-tuning of Large Language Models (LLMs) using Low-Rank Adaptation (LoRA) on an end-user's data offers personalized experiences while keeping data private, but faces severe memory constraints on consumer hardware.
By Hassan Dbouk, Matthias Reisser, Prathamesh Mandke, Likhita Arun Navali, Christos Louizos