arXiv Computation and Language By Chandra Vamsi Krishna Alla, Harish Naidu Gaddam, Manohar Kommi, Sheikh Nazib Ahmed

BudgetMem: Training-Free Selective Memory for Cost-Efficient Long-Context Processing in Language Models

Read the original on arXiv Computation and Language →

arXiv:2511. 04919v3 Announce Type: replace Abstract: Processing long documents with large language models (LLMs) is expensive: a single query over a 100K-token document can cost from tens of cents to over a dollar in API fees, depending on the model, and memory grows linearly with context length.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 24

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

arXiv:2606. 24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware.

By Xiaolin Lin, Jingcun Wang, Olga Kondrateva, Yiyu Shi, Bing Li, Grace Li Zhang