← Back to all news
arXiv Machine Learning September 1, 2026 By Tong Yuan, Chengxi Liao, Zeyi Wen

Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • agents
  • nlp
  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Aug 31

Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

Long-context LLM applications such as document summarization and multi-turn agents require generation from prefixes spanning tens of thousands of tokens, making decoding latency a major bottleneck. Sp...

llmsagentsnlpefficiency
More like this →
arXiv AI
Jun 2

BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding

arXiv:2606. 00144v1 Announce Type: cross Abstract: Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel.

By Liang He, Jingbo Wen, Qishi Zhan, Yixiong Chen, Kangning Cui, Qizhen Lan, Xilu Wang
efficiency
More like this →
Hugging Face Trending Papers
Jul 23

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap.

efficiency
More like this →
arXiv Machine Learning
Jul 24

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

arXiv:2607. 21535v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel.

By Alagappan Valliappan
efficiency
More like this →
arXiv Computation and Language
Aug 27

AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs

arXiv:2608. 26004v1 Announce Type: cross Abstract: Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions.

By Sheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv, Hao Wang, Chen Zhang, Yong Liu
llmsagentsefficiencybenchmarks
More like this →
arXiv AI
Jun 16

Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving

arXiv:2512. 22420v5 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel.

By Rui Li, Zhaoning Zhang, Libo Zhang, Huaimin Wang, Xiang Fu, Zhiquan Lai
llmsefficiencybenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea