← Back to all news
Hugging Face Blog November 20, 2024

Faster Text Generation with Self-Speculative Decoding

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Jul 9

A Practical Investigation of Training-free Relaxed Speculative Decoding

Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verified in parallel by the LLM. Standard speculative decoding is lossless: its rejection and resampling steps exactly preserve the LLM's sampling distribution.

llmsbenchmarks
More like this →
arXiv AI
Jul 10

A Practical Investigation of Training-free Relaxed Speculative Decoding

arXiv:2607. 08690v1 Announce Type: cross Abstract: Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verified in parallel by the LLM.

By Guoxuan Xia, Luka Ribar, Paul Balanca
llmsbenchmarks
More like this →
Hugging Face Blog
May 29, 2024

Benchmarking Text Generation Inference

benchmarks
More like this →
arXiv AI
Jun 4

SSSD: Simply-Scalable Speculative Decoding

arXiv:2411. 05894v3 Announce Type: replace-cross Abstract: Speculative Decoding has emerged as a popular technique for accelerating inference in Large Language Models.

By Michele Marzollo, Jiawei Zhuang, Niklas Roemer, Niklas Zwingenberger, Lorenz K. M\"uller, Lukas Cavigelli
llmsbenchmarks
More like this →
Hugging Face Trending Papers
Jul 12

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks. However, traditional speculative decoding typically relies on auxiliary draft modules, incurring significant training and communication overhead.

llmsefficiencybenchmarks
More like this →
Hugging Face Blog
Oct 8, 2024

Faster Assisted Generation with Dynamic Speculation

More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e