Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
Read the original on arXiv Computation and Language →Large Language Models suffer from high inference latency, especially when generating long chains of thought. Existing speculative decoding methods draft and verify tokens in parallel but ignore semantic equivalence, causing inefficient rejections. The proposed SemanticSpec framework verifies entire semantic sequences by probing internal hidden states, achieving up to 2.7× speedup on DeepSeekR1-32B and 2.1× on QwQ-32B while outperforming token‑level and sequence‑level baselines in both efficiency and effectiveness.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.