Hugging Face Trending Papers

AdaPLD: Adaptive Retrieval and Reuse for Efficient Model-Free Speculative Decoding

Read the original on Hugging Face Trending Papers →

Speculative decoding accelerates generation by verifying multiple drafted tokens in a single target-model forward pass, reducing sequential decoding iterations. Model-free variants avoid auxiliary draft models by reusing text and model states already available during generation, but their speedup depends on the reliability of the constructed drafts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

Hugging Face Trending Papers
Sep 17

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

The paper introduces SwitchSD, an adaptive framework that treats copying as a latent control signal in large language model decoding. By training lightweight probes on internal representations, SwitchSD accurately detects genuine copy intent (AUC > 0.99) and dynamically switches between neural drafting and context-based copying. Experiments on Llama and Qwen models show up to 15 % throughput gains over state‑of‑the‑art baselines such as EAGLE3.

arXiv Computation and Language
Sep 18

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

The paper introduces SwitchSD, an adaptive framework that controls speculative decoding by distinguishing genuine copy intent from accidental repetitions using lightweight probes on a model’s internal representations. SwitchSD dynamically switches between neural drafting and context-based copying, achieving up to 15% throughput gains over existing baselines such as EAGLE3. The approach turns copying from a noisy heuristic into a principled, model-aware decoding regime.

By Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik, Lior Wolf, Itamar Zimerman
Hugging Face Trending Papers
Jul 21

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass.