← Back to all news
Hugging Face Trending Papers August 31, 2026

Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • agents
  • nlp
  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Sep 1

Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

arXiv:2608.30252v1 Announce Type: new Abstract: Long-context LLM applications such as document summarization and multi-turn agents require generation from prefixes spanning tens of thousands of token...

By Tong Yuan, Chengxi Liao, Zeyi Wen
llmsagentsnlpefficiency
More like this →
arXiv Machine Learning
1d ago

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

arXiv:2609.17943v1 Announce Type: new Abstract: Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates...

By Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao, Murali Annavaram, Salman Avestimehr
llmsefficiencybenchmarks
More like this →
arXiv Computation and Language
Aug 27

AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs

arXiv:2608. 26004v1 Announce Type: cross Abstract: Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions.

By Sheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv, Hao Wang, Chen Zhang, Yong Liu
llmsagentsefficiencybenchmarks
More like this →
arXiv AI
Jun 16

Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving

arXiv:2512. 22420v5 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel.

By Rui Li, Zhaoning Zhang, Libo Zhang, Huaimin Wang, Xiang Fu, Zhiquan Lai
llmsefficiencybenchmarks
More like this →
Hugging Face Trending Papers
Jul 23

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap.

efficiency
More like this →
arXiv AI
Jun 2

Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding

arXiv:2606. 01019v1 Announce Type: cross Abstract: Large Language Model (LLM) generation remains expensive because autoregressive decoding calls the model once for each new token.

By Xin Su, Dawid Majchrowski, Fangyuan Yu, Vanshil Atul Shah, Sebastian Rogawski, Pawel Morkisz, Anahita Bhiwandiwalla, Phillip Howard
llmsagents
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea