arXiv:2605. 03644v2 Announce Type: replace Abstract: Many-Shot In-Context Learning (ICL) has emerged as a promising paradigm, leveraging extensive examples to unlock the reasoning potential of Large Language Models (LLMs).
By Jie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Ruiqi Wu, Zhaokun Wang, Wenyi Li, Wenhong Tian
arXiv:2609.25537v1 Announce Type: new
Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing laten...
By Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia, Yutaka Watanobe, Fang Liu
The quadratic growth of attention computation and key-value (KV) cache with respect to sequence length is a central bottleneck for ultra-long-context language models and high-resolution generative mod...
arXiv:2609.37988v1 Announce Type: new
Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This i...
By Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi
arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.
By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt.
The paper introduces JoLT, a training‑free compressor that jointly allocates rank and precision for key‑value (KV) cache compression in long‑context language models. JoLT treats grouped prefill caches as fourth‑order tensors, applies partial Tucker decomposition along token and feature modes, and uses a rotated low‑bit quantizer for residuals, all governed by a single Lagrangian dual under a global byte constraint. Across five models from four architecture families, JoLT achieves 2–3× compression with less than 0.2% perplexity loss, and near‑lossless retrieval accuracy on LLaMA‑3.1‑8B at 64K context up to 3× compression.
By Rahul Krishnan, Volker Schulz
arXiv:2606. 11853v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by finite context windows and the growing cost of key-value (KV) caches in long multi-modal sequences.
By Zhirui Chen, Ziwei Chen, Ling Shao
VisCache introduces a two-stage, plug‑and‑play framework for pruning visual key‑value caches in Vision Large Language Models without retraining. The first stage filters out temporally redundant keyframes, while the second stage, PruneKV, applies a parabolic layer‑wise budget and asymmetric update to selectively prune keys and fuse values, preserving essential context. Experiments show up to 2.35× speedup and significant memory savings with only 19–28% of the original cache retained, outperforming existing baselines.
By Lyuke Wang, Zhuo Li, Guangxu Zhu
arXiv:2608.23843v1 Announce Type: new
Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compressio...
By Zizhong Wang, Jieying Wang, Zhao Zhang, Jiajia Li
The paper investigates KV cache offloading for long‑context large language models, focusing on tasks that require extensive information extraction from the prompt. The authors introduce the Text2JSON benchmark, a highly context‑intensive task that demands structured knowledge extraction from raw text, and evaluate modern KV offloading techniques on this benchmark and other similar tasks. Their experiments on Llama 3 and Qwen 3 reveal significant accuracy degradation, attributing it to low‑rank key projection and unreliable landmarks, and propose a simpler strategy that markedly improves performance across multiple LLM families.
By Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy, Yegor Yershov
arXiv:2607. 14327v1 Announce Type: cross Abstract: Efficient long-context inference is not only about reducing memory cost, but also about keeping useful contextual evidence accessible as generation proceeds.
By Bohan Yu, Lei Shen, Chenxi Zhou, Chen Han, Junlin Liu, Wenbo Su, Yu Cheng, Bo Zheng