Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
Read the original on arXiv AI →arXiv:2607. 17283v1 Announce Type: new Abstract: Single-stream autoregressive decoding of large language models is bound by memory bandwidth: each generated token requires one full forward pass through the target model, and successive passes cannot be parallelized.
Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.