arXiv:2607. 14107v1 Announce Type: cross Abstract: The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-cache reuse, while increasing decoding parallelism with static confidence thresholds can compromise generation quality.
By Mingyu Lee, Akshat Ramachandran, Souvik Kundu, Tushar Krishna
arXiv:2608. 03276v1 Announce Type: new Abstract: Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length.
By Wonpyo Park, Seung-won Hwang
arXiv:2609.39329v1 Announce Type: new
Abstract: Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compressi...
By Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi, Donggyun Kim, Seunghoon Hong
Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt.
arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.
By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
SGD-KV is a head‑aware framework for compressing key‑value caches in large language models. It uses a chunk‑summarization diagnostic task to identify attention heads that specialize in hierarchical information aggregation, allowing the KV cache budget to be allocated based on each head’s summarization score. Experiments on Qwen2.5‑7B‑1M and Qwen3‑32B show state‑of‑the‑art performance on up to 1M‑token contexts while cutting KV cache memory usage by up to 75%.
By Zeyu Liu, Woomin Song, Xuandi Fu, Sai Muralidhar Jayanthi, Vivek Govindan, Aram Galstyan, Sravan Babu Bodapati, Srikanth Ronanki
arXiv:2607. 06523v1 Announce Type: new Abstract: Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply uniform budgets across layers or tokens and degrade retrieval when lexical cues and semantic states require different preservation.
By Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jesus Olivera
The paper introduces ACache, an affix-oriented cache reuse mechanism for Diffusion Large Language Models (DLLMs). ACache identifies a small set of critical affix tokens, called Anchor Tokens, and selectively recomputes their key-value states while reusing the rest of the affix cache. Experiments on Fast-dLLM and Nano-vLLM show that recomputing about 20% of affix tokens restores accuracy and can reduce recompute latency by up to 55.7% while improving throughput by up to 1.68×.
By Kaihua Liang, An Zhong, Xin Tan, Zafar Ayyub Qazi, Hong Xu, Jian Weng, Marco Canini
SGD-KV is a head‑aware framework that compresses key‑value caches in large language models by using a chunk‑summarization diagnostic task to identify attention heads that specialize in hierarchical information aggregation. It prioritizes these heads during compression, achieving state‑of‑the‑art performance on long‑context benchmarks with up to 1M tokens while cutting KV cache memory usage by as much as 75%. Experiments on Qwen2.5‑7B‑1M and Qwen3‑32B confirm that allocating cache budget based on summarization scores yields a superior efficiency‑accuracy trade‑off for long‑context inference.
arXiv:2608.30295v1 Announce Type: cross
Abstract: Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challeng...
By Haoyun Jiang, Haolin Li, Jianwei Zhang, Fei Huang, Qiang Hu, Minmin Sun, Shuai Xiao, Yong Li, Junyang Lin, Jiangchao Yao
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
By Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin
arXiv:2605. 25475v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows linearly with sequence length, quickly becoming the bottleneck for long context inference.
By Xintong Yang, Hao Gu, Binxing Xu, Lujun Li, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Yike Guo, Sirui Han