arXiv Computation and Language By Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik, Lior Wolf, Itamar Zimerman

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

Read the original on arXiv Computation and Language →

The paper introduces SwitchSD, an adaptive framework that controls speculative decoding by distinguishing genuine copy intent from accidental repetitions using lightweight probes on a model’s internal representations. SwitchSD dynamically switches between neural drafting and context-based copying, achieving up to 15% throughput gains over existing baselines such as EAGLE3. The approach turns copying from a noisy heuristic into a principled, model-aware decoding regime.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Sep 17

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

The paper introduces SwitchSD, an adaptive framework that treats copying as a latent control signal in large language model decoding. By training lightweight probes on internal representations, SwitchSD accurately detects genuine copy intent (AUC > 0.99) and dynamically switches between neural drafting and context-based copying. Experiments on Llama and Qwen models show up to 15 % throughput gains over state‑of‑the‑art baselines such as EAGLE3.